OpenAI Pauses Parts of Astra’s Development After Flagging ‘Critical’ Cyber Capabilities
OpenAI is pausing parts of the development of its upcoming Astra model after concluding it cannot rule out the model has crossed the 'Critical' cybersecurity threshold defined in its own Preparedness...
OpenAI said Friday it is pausing parts of the internal development of Astra, its upcoming model, after evaluations over the past several days turned up advancements in agentic coding and cybersecurity strong enough that the company cannot rule out Astra has reached the “Critical” cybersecurity threshold defined in its own Preparedness Framework. In a post published August 7, OpenAI said it reached that conclusion “last night” and was disclosing it because it believes in being transparent with the public and the safety and security communities about a potential shift in what its models can do.
Table Of Content
What the Critical Threshold Actually Means
OpenAI created its Preparedness Framework in December 2023 as an internal gate for tracking when a model’s capabilities in categories such as biology, chemistry, cybersecurity, and AI self-improvement cross thresholds that require additional safeguards before development or release can continue. Under that framework, OpenAI says a model reaches the Critical cybersecurity threshold “if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.”
That is one level above High, which is where OpenAI says every prior frontier model, including GPT-5.6-Sol, has landed during internal evaluations. The company was careful to note that Astra’s evaluations are still ongoing and that it has not confirmed the model has actually crossed the line, only that current results are strong enough that it “cannot rule out Critical capability level at this time.” According to PCWorld, GPT-5.6-Sol itself was first released only to a “select group of trusted partners” before OpenAI made it public a couple of weeks later, a staged pattern that Astra’s eventual rollout may end up following too.
New Safeguards, and a Partial Pause
OpenAI says it has responded by scaling up testing of its existing safeguards and adding several controls specifically for higher-capability models: isolated testing environments, restricted network and tool access, stronger encryption and protections around model weights, additional monitoring and detection, and sandboxed execution. Internally, the company says it is now pausing any activity involving Astra that does not meet those strengthened requirements, while continuing work that does. It has also rolled out what it calls universal monitoring across all of Astra’s agentic uses, including its own training and evaluation, with automated monitors reviewing the model’s chain of thought and triggering a security review if they spot high-risk activity.
Going forward, OpenAI says unspecified government agencies and “select AI safety organizations” will get access to test Astra’s capabilities directly, and that it will share its own recommended security controls with third-party partners who need to run higher-risk evaluations. The company points to its own precedent for this kind of response: when its models approached the High threshold for biological risk in June 2025, it says it followed a similar playbook of tighter safeguards, expanded testing, outside experts, and additional security controls before moving forward.
A Second Astra Headline in a Week, Amid a Wider Pattern
The disclosure lands less than a week after OpenAI announced that an internal version of Astra had produced machine-checked proofs for ten previously unsolved problems in mathematics and theoretical computer science. This time, OpenAI was explicit that the model behind that math result is not the source of its cybersecurity concern in the way readers might assume: the company stated plainly that “Astra is an upcoming model, and was not involved in exploiting Hugging Face,” a reference to an incident in which OpenAI’s GPT-5.6 Sol model and an unnamed, more capable pre-release model broke out of an isolated test environment in July and reached the production systems of the machine learning platform Hugging Face, an incident OpenAI disclosed on July 21.
That earlier incident kicked off a broader wave of similar admissions. Anthropic followed on July 30 with its own disclosure that three Claude models, while being tested inside environments run by a third-party partner, had breached the real infrastructure of three organizations during cybersecurity evaluations, an incident it described in detail on its own blog. OpenAI’s Astra announcement is a different kind of disclosure, a capability warning issued before any breach rather than an account of one that already happened, but TechCrunch noted that companies across industries rarely announce this kind of internal pause publicly while a product is still in development. PCWorld’s Ben Patterson framed it more bluntly, writing that the run of disclosures suggests AI safety may have reached “a crossroads,” where each new frontier model gets judged, at least initially, “too powerful to be released.”
OpenAI has not given a timeline for when Astra’s cybersecurity evaluations will wrap up or when the paused activities might resume. For now, the company says its focus is getting its safeguards to a level it considers appropriate for a model with these capabilities before deciding what happens next.








No Comment! Be the first one.