OpenAI is slowing work on its upcoming artificial intelligence model, Astra, after internal evaluations identified significant advances in its agentic coding and cybersecurity capabilities, prompting the company to introduce stricter safeguards before deployment.
The AI company said it cannot rule out the possibility that Astra could reach the “Critical” cybersecurity capability threshold under its Preparedness Framework, a classification that triggers additional security measures for models capable of potentially causing severe harm.
Astra has not been formally released, and OpenAI’s decision does not amount to a complete halt in its development. Instead, the company is pausing activities involving the model that do not meet its strengthened security requirements while additional protections are implemented.
Under OpenAI’s framework, a model could reach the Critical threshold if it is capable of autonomously identifying and developing functional zero-day exploits across hardened, real-world critical systems. The classification also covers the ability to devise and execute novel end-to-end cyberattack strategies without human intervention.
OpenAI said its latest evaluations showed significant improvements in agentic coding and cybersecurity. These capabilities could have legitimate applications, including identifying software vulnerabilities and supporting defensive security work, but they could also create risks if used to automate cyberattacks or vulnerability exploitation at scale.
The company is now introducing additional safeguards around Astra. Planned measures include using isolated testing environments, restricting network and tool access, implementing sandboxed execution and expanding monitoring of how the model operates.
OpenAI also plans to work with government agencies, AI safety institutes and other organisations to evaluate the model before wider deployment.
The decision highlights a growing challenge for AI developers as frontier models become more capable of performing complex tasks with less human involvement. Agentic AI systems are designed to carry out multi-step actions and use external tools, making their ability to operate autonomously an increasingly important part of safety evaluations.
Cybersecurity has emerged as a particularly sensitive area because advanced AI models can potentially be used for both defensive and offensive purposes. The same capabilities that allow a system to identify vulnerabilities, analyse code and assist security professionals could also be misused to find weaknesses or automate elements of an attack.
OpenAI’s move comes as other AI companies face similar questions about managing increasingly capable systems. Anthropic, Meta and other developers have been strengthening testing and safeguards around models capable of coding, tool use and autonomous actions.
Astra’s eventual release timeline remains unclear. The model has not been formally announced as a commercial product, and OpenAI has not provided a confirmed date for when it could become publicly available.
The development also shows how safety frameworks are beginning to influence the pace at which frontier AI systems move from internal testing to broader deployment. As companies race to develop models with stronger reasoning, coding and autonomous capabilities, evaluations are increasingly being used to determine whether additional controls are required before release.
For OpenAI, Astra represents a test of that process. Development can continue under tighter controls, but broader deployment will depend on whether the company can establish safeguards appropriate to the cybersecurity capabilities identified during testing.