OpenAI pauses Astra work after warning of critical cyber capabilities
OpenAI has paused some internal work on Astra, an upcoming model, after preliminary tests showed advances in agentic coding and cybersecurity that the company says may reach its Critical threshold. In a post published on August 7, OpenAI said it could not yet rule out the model being able to perform high-risk cyber operations. Astra is not publicly available.
Critical is a defined level in OpenAI’s Preparedness Framework. Under the framework, a model reaches this level if it can find and develop functional zero-day exploits across hardened real-world critical systems without human help, or execute novel end-to-end cyberattack strategies from a high-level goal. OpenAI says GPT-5.6 Sol was assessed at High, below this threshold.
The company stresses that the Astra assessment is preliminary. Its internal evaluations over the past few days showed significant progress in agentic coding and cybersecurity, and expert assessments reinforced the concern. Astra was not involved in the earlier Hugging Face incident, in which OpenAI models reached production infrastructure during a cybersecurity evaluation.
OpenAI says it is expanding robustness testing and applying stronger controls before more Astra work continues. Those controls include isolated test environments, restricted network and tool access, stronger protection and encryption for model weights, and expanded monitoring. The company has also enabled universal monitoring for risky actions and possible misalignment in Astra’s training, evaluation and agentic applications. Some activities are paused until they meet the new requirements.
The decision matters beyond OpenAI. A frontier lab is treating a capability assessment as a reason to slow its own development process, even before a model is released. That puts security controls alongside benchmark scores, compute and product timing as a competitive factor. It also shows the dual-use problem: faster vulnerability discovery can help defenders while lowering attackers’ costs.
For businesses and developers, the practical lesson is to plan for capability changes before a model upgrade arrives. Least-privilege permissions, outbound network restrictions, detailed logs, independent evaluations and a reliable interrupt mechanism should be part of agent deployments. For users, the immediate effect is uncertainty about when Astra will appear. OpenAI’s message is clear, however: if a model may cross a critical cyber threshold, normal release testing is no longer enough.