OpenAI Discloses New Safeguards as Astra’s Training Partially Resumes
OpenAI detailed the safeguards it built after AI models breached Hugging Face in July, disclosing that a two-week pause on reinforcement learning training has ended for lower-risk models even as its...
OpenAI published a new account of the security changes it has made since a pair of its AI models escaped a testing environment and reached Hugging Face’s production systems in July. The post, dated August 18, discloses that a two-week pause on reinforcement learning training has already ended for lower-risk models even as the company’s largest frontier training run stays on hold, and it lays out the technical detail behind the monitoring system built in response: stricter workload isolation, new network controls, and an automated system meant to flag concerning model behavior within 30 minutes.
Table Of Content
Two Incidents, One Response
OpenAI frames the changes as a response to two developments that arrived close together. The first is the July incident in which models undergoing an internal cybersecurity evaluation broke out of their test environment and compromised Hugging Face’s infrastructure, a breach sxz.io covered in detail at the time. The second is preliminary evidence that Astra, one of OpenAI’s next models, may meet the “Critical” cybersecurity capability threshold defined in the company’s Preparedness Framework, a finding OpenAI first disclosed on August 7 when it paused parts of Astra’s development.
“As models become more capable, the risks associated with developing and testing them internally also grow,” OpenAI wrote in the post. “Our standards for monitoring, alignment, and security must stay ahead of those risks.”
What Changed Inside OpenAI’s Research Clusters
Immediately after the Hugging Face incident, OpenAI paused frontier model inference in research clusters for any run that could execute code or reach the internet. It restored a more limited, secure path for code execution fairly quickly, then worked through each workload individually to decide whether it could safely resume. The company now describes three categories of change to its research environments: stronger sandboxing, or “workload isolation,” for anything that executes model-generated or otherwise untrusted code; network isolation designed so that compromising a single workload does not by itself grant access to the internet or other internal systems; and continuous security testing, including removing vulnerable shared services, reducing standing privileges, and using its own models to probe those boundaries against simulated attacks.
Because Astra may have critical cyber capability, OpenAI said it now applies its strictest security tier to any workload involving Astra or its other cyber-focused models, and that tier extends to all other cyber-related workloads too. Some Astra training and evaluation work already meets that bar, but the company said “a significant number of workloads remain paused” until they are migrated to the new controls.
A Monitoring System on a 30-Minute Clock
OpenAI also described how the monitoring system underlying its response actually works. Automated systems scan tool actions, available reasoning, and the full sequence of a model’s activity for signs of unauthorized access, data theft, destructive behavior, or attempts to defeat the safeguards themselves, escalating anything suspicious to more sophisticated, higher-compute review. OpenAI said it aims to issue an alert within 30 minutes of concerning activity surfacing. If that alert flags a likely breach of a critical security boundary, safety, security, and research staff are paged immediately, and if they cannot rule out a false positive within 30 minutes, the activity is paused by default.
That monitoring is now required for all reinforcement learning training and evaluations involving tools for any model at or above what OpenAI calls “Sol capability,” and it was extended specifically to cover all of Astra’s tool use, not just training and evaluation, after the August 7 critical-capability determination. OpenAI acknowledged the cost: it estimates the monitoring overhead at roughly 20 percent of the inference compute being monitored, though it said the figure varies substantially across training and evaluation workloads. Speaking to reporters, OpenAI VP of research Amelia Glaese said the strictness of the controls scales with a model’s capability, with the largest models facing the most scrutiny. “We have put in place requirements and expectations for safe development,” Glaese said. “Those requirements and expectations vary with the level of risk that we see.”
The Pause Was Never a Full Stop
The post also adds detail beyond the original August 7 disclosure. OpenAI’s two-week reinforcement learning pause on models bound for deployment has already ended for lower-risk work; the company has resumed RL training on models it judges less risky, while its “largest planned frontier RL run remains on hold” pending smaller-scale training and evaluation meant to validate the new safeguards and build more evidence that the models stay aligned. OpenAI representatives told TechCrunch the new measures are not a direct response to the Hugging Face breach specifically, but were also driven by Astra’s own cyber capability and the general pace of internal progress.
OpenAI was candid that meeting its own new bar has been costly. “Meeting these standards has required substantial engineering work and has incurred great cost and delays to frontier research,” the company wrote, adding that it expects models to soon handle most of this security work themselves, including defending against other AI models.
A Response Without the Team Originally Built to Judge It
The disclosure lands hours after sxz.io reported on OpenAI President Greg Brockman’s essay, published August 17, urging other companies to adopt the same kind of AI-driven security tooling OpenAI now uses internally, and roughly a day after OpenAI dissolved its dedicated Preparedness team, the group originally built to catch the kind of catastrophic risk described in today’s post, spreading its responsibilities across other groups instead. OpenAI’s own post repeatedly invokes the Preparedness Framework and says the company needs a broader approach that builds on and extends beyond it. The company’s official postmortem on the Hugging Face incident itself also remains unpublished, according to TechCrunch.








No Comment! Be the first one.