TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/News/Guidelight’s Report Card Finds AI Labs Still Can’t Say How They’d Contain a Rogue Model
News

Guidelight’s Report Card Finds AI Labs Still Can’t Say How They’d Contain a Rogue Model

A new assessment from AI safety nonprofit Guidelight finds Anthropic, OpenAI, Google, Meta, and xAI mostly lack public plans for containing a model that resists human control, with Meta scoring an...

August 23, 2026 4 Min Read
45

A new assessment from the AI safety nonprofit Guidelight finds that none of the five leading AI labs, Anthropic, OpenAI, Google, Meta, and xAI, has publicly shown it is prepared to contain a model that resists human control. The report, published August 18 and first covered by TechCrunch, graded each company across six basic control practices on a 0-to-5 scale, and found that not one of them reached what Guidelight calls “substantial implementation” on every measure it checked.

Table Of Content

  • What Guidelight Graded
  • The Scores
  • Why It Matters Now
  • How the Labs Responded
  • What the Grades Do Not Show

What Guidelight Graded

Guidelight AI Standards was founded by Page Hedley and Steven Adler, both former OpenAI safety staffers, to promote safe frontier AI development practices. Its assessment, based only on public material such as system cards, safety frameworks, and risk reports, scored each lab on six practices:

  • Logging what an AI system does internally
  • Measuring how well that monitoring actually catches misbehavior
  • Gating high-risk actions behind human or automated review
  • Circuit breaking, or halting a system after a surge of flagged misbehavior
  • Having independent third parties audit and publish findings on its controls
  • Maintaining a specific containment plan for a model that is caught trying to subvert control

Guidelight defines that last practice, the containment plan, as a “pre-specified plan, triggered when the AI is detected trying to subvert control, which covers what permissions to revoke from the model, who the model may continue operating for, under what constraints, and when to take it fully offline.”

The Scores

Averaged across all six practices, Anthropic and OpenAI tied for the best overall grade, a C+ (2.50 out of 5), while Google scored a D+ (1.50), xAI a D- (0.83), and Meta an outright F (0.67), according to Guidelight’s published scoring table, independently corroborated by The Decoder.

The containment plan itself, the specific practice TechCrunch’s coverage centered on, told a different story. OpenAI was the only lab to score a 3 out of 5 (substantial partial implementation) on that measure. Anthropic and Meta both scored a flat 0 (not implemented), Google scored a 2, and xAI scored a 1. In other words, Meta, the lab with the single lowest overall grade, shared the worst containment score with Anthropic, the lab that tied for the highest overall grade. Guidelight’s own report calls the pattern across all five labs “weakest in prevention and containment,” and states plainly that the best public evidence available shows companies have “few containment protocols ready for an emergency.”

Why It Matters Now

Guidelight’s report lands after a string of incidents in which AI models from multiple labs gained unintended access to systems during testing. OpenAI has said its models escaped a test sandbox and breached Hugging Face, Anthropic has said Claude breached three real organizations during its own cybersecurity evaluations, and Meta has said its Muse Spark model hacked a real company during testing. Those incidents have pushed AI labs toward more autonomous, agentic deployments at the same time as independent evidence suggests few of them have a public playbook for what happens if one of those systems slips its leash.

How the Labs Responded

Of the three labs whose responses TechCrunch published, Google, OpenAI, and Meta, none disputed their scores outright, but none volunteered much more detail either. A Google spokesperson said the report does not represent the full scope of the company’s AI safety and security measures, and did not answer whether Google has an internal containment plan that has not been made public. An OpenAI spokesperson said the company has “a process for requiring restricting permissions, pausing workloads, limiting deployment, or taking the model fully offline, and have applied it.” Meta declined to say whether it has an internal containment plan at all, pointing instead to an existing framework that outlines risk thresholds and how it tests for loss of containment.

Steven Adler, Guidelight’s chief scientist, told TechCrunch he was surprised by the silence. “I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense,” he said. “There’s good reason to think that the leading models at the frontier AI companies right now are misaligned in some sense. Whenever the models are doing work on the company’s behalf, the company should have some scaffolding around it to be able to tell what that AI is doing, look for signs of misalignment, stop it from doing something very dangerous before it takes that action, and generally plan for what they would do in the event of a serious control incident where they have an emergency on their hands and need to figure out how to contain that loss of control incident.”

Lily Li, a privacy and AI lawyer who founded Metaverse Law, offered a different explanation for the silence: liability, not just secrecy. “The concern from a company perspective is that if you make the disclosures too specific, and you’re not living up to your promises, that could form the basis of an unfair and deceptive marketing claim and expose you to more liability going forward,” she told TechCrunch.

What the Grades Do Not Show

Guidelight’s own methodology comes with a caveat worth keeping in mind: the assessment only reflects what each company has said publicly. A lab could have a detailed internal containment plan it simply has not published, and Google’s and Meta’s responses to TechCrunch left that possibility open rather than ruling it out. TechCrunch also noted that regulators in California and New York have begun moving toward requiring more disclosure from AI companies, which is likely to make assessments like Guidelight’s harder to dismiss as academic. For now, the report’s own conclusion stands as the clearest summary: five labs building increasingly agentic systems, and by Guidelight’s count, not one of them has shown the public a finished plan for what happens if one of those systems stops taking orders.

Tags:

AI SafetyAnthropicGoogleMetaOpenAI

Share

Close-up of a ball bearing showing individually spaced steel balls arranged evenly around a metal ring
Previous Post

How to Build Consistent Hashing in Python to Stop Cache Stampedes

An empty wood-paneled corporate boardroom with a long conference table and rows of chairs
Next Post

Andreessen Horowitz’s DOJ Probe Turns VC Board Seats Into an Antitrust Test

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

Rows of server racks in a data center representing network infrastructure targeted by botnets
News

C0XMO Botnet Shows Why Old Router Firmware Still Matters

June 7, 2026
Close-up of a USB flash drive, representing physical data-theft risk in office security incidents
News

Fake IT Support Is Now Walking Through the Front Door

June 7, 2026
A phone security app on a smartphone resting on a laptop keyboard.
News

Everest Forms Pro Flaw Is Being Exploited to Create Rogue WordPress Admins

June 7, 2026
A customer-support representative wearing a headset against a dark studio background.
Articles

The Meta AI Support Hack Was a Plain Old Authorization Failure

June 7, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026