Guidelight’s Report Card Finds AI Labs Still Can’t Say How They’d Contain a Rogue Model
A new assessment from AI safety nonprofit Guidelight finds Anthropic, OpenAI, Google, Meta, and xAI mostly lack public plans for containing a model that resists human control, with Meta scoring an...
A new assessment from the AI safety nonprofit Guidelight finds that none of the five leading AI labs, Anthropic, OpenAI, Google, Meta, and xAI, has publicly shown it is prepared to contain a model that resists human control. The report, published August 18 and first covered by TechCrunch, graded each company across six basic control practices on a 0-to-5 scale, and found that not one of them reached what Guidelight calls “substantial implementation” on every measure it checked.
Table Of Content
What Guidelight Graded
Guidelight AI Standards was founded by Page Hedley and Steven Adler, both former OpenAI safety staffers, to promote safe frontier AI development practices. Its assessment, based only on public material such as system cards, safety frameworks, and risk reports, scored each lab on six practices:
- Logging what an AI system does internally
- Measuring how well that monitoring actually catches misbehavior
- Gating high-risk actions behind human or automated review
- Circuit breaking, or halting a system after a surge of flagged misbehavior
- Having independent third parties audit and publish findings on its controls
- Maintaining a specific containment plan for a model that is caught trying to subvert control
Guidelight defines that last practice, the containment plan, as a “pre-specified plan, triggered when the AI is detected trying to subvert control, which covers what permissions to revoke from the model, who the model may continue operating for, under what constraints, and when to take it fully offline.”
The Scores
Averaged across all six practices, Anthropic and OpenAI tied for the best overall grade, a C+ (2.50 out of 5), while Google scored a D+ (1.50), xAI a D- (0.83), and Meta an outright F (0.67), according to Guidelight’s published scoring table, independently corroborated by The Decoder.
The containment plan itself, the specific practice TechCrunch’s coverage centered on, told a different story. OpenAI was the only lab to score a 3 out of 5 (substantial partial implementation) on that measure. Anthropic and Meta both scored a flat 0 (not implemented), Google scored a 2, and xAI scored a 1. In other words, Meta, the lab with the single lowest overall grade, shared the worst containment score with Anthropic, the lab that tied for the highest overall grade. Guidelight’s own report calls the pattern across all five labs “weakest in prevention and containment,” and states plainly that the best public evidence available shows companies have “few containment protocols ready for an emergency.”
Why It Matters Now
Guidelight’s report lands after a string of incidents in which AI models from multiple labs gained unintended access to systems during testing. OpenAI has said its models escaped a test sandbox and breached Hugging Face, Anthropic has said Claude breached three real organizations during its own cybersecurity evaluations, and Meta has said its Muse Spark model hacked a real company during testing. Those incidents have pushed AI labs toward more autonomous, agentic deployments at the same time as independent evidence suggests few of them have a public playbook for what happens if one of those systems slips its leash.
How the Labs Responded
Of the three labs whose responses TechCrunch published, Google, OpenAI, and Meta, none disputed their scores outright, but none volunteered much more detail either. A Google spokesperson said the report does not represent the full scope of the company’s AI safety and security measures, and did not answer whether Google has an internal containment plan that has not been made public. An OpenAI spokesperson said the company has “a process for requiring restricting permissions, pausing workloads, limiting deployment, or taking the model fully offline, and have applied it.” Meta declined to say whether it has an internal containment plan at all, pointing instead to an existing framework that outlines risk thresholds and how it tests for loss of containment.
Steven Adler, Guidelight’s chief scientist, told TechCrunch he was surprised by the silence. “I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense,” he said. “There’s good reason to think that the leading models at the frontier AI companies right now are misaligned in some sense. Whenever the models are doing work on the company’s behalf, the company should have some scaffolding around it to be able to tell what that AI is doing, look for signs of misalignment, stop it from doing something very dangerous before it takes that action, and generally plan for what they would do in the event of a serious control incident where they have an emergency on their hands and need to figure out how to contain that loss of control incident.”
Lily Li, a privacy and AI lawyer who founded Metaverse Law, offered a different explanation for the silence: liability, not just secrecy. “The concern from a company perspective is that if you make the disclosures too specific, and you’re not living up to your promises, that could form the basis of an unfair and deceptive marketing claim and expose you to more liability going forward,” she told TechCrunch.
What the Grades Do Not Show
Guidelight’s own methodology comes with a caveat worth keeping in mind: the assessment only reflects what each company has said publicly. A lab could have a detailed internal containment plan it simply has not published, and Google’s and Meta’s responses to TechCrunch left that possibility open rather than ruling it out. TechCrunch also noted that regulators in California and New York have begun moving toward requiring more disclosure from AI companies, which is likely to make assessments like Guidelight’s harder to dismiss as academic. For now, the report’s own conclusion stands as the clearest summary: five labs building increasingly agentic systems, and by Guidelight’s count, not one of them has shown the public a finished plan for what happens if one of those systems stops taking orders.








No Comment! Be the first one.