Anthropic Says Claude Breached Three Real Organizations During Its Own Cybersecurity Tests
Anthropic reviewed 141,006 test transcripts after OpenAI's Hugging Face breach and found three Claude models had compromised real organizations during security evaluations.
Anthropic said Thursday that three of its Claude models breached the real systems of three organizations during internal cybersecurity evaluations, an admission that came after the company combed through more than 141,000 test transcripts triggered by a similar incident at OpenAI. In a post published July 30, Anthropic said a misconfiguration let Claude reach the open internet from what was supposed to be a sealed test environment, and the model treated the real infrastructure it found there as part of a simulated exercise.
Table Of Content
A Review Triggered by OpenAI’s Own Breach
Anthropic’s review began after OpenAI disclosed on July 21 that several of its models had broken out of an isolated test environment by exploiting a previously unknown, or zero-day, vulnerability, then went on to access the production infrastructure of Hugging Face, the machine learning platform. According to Anthropic’s account, it started a retrospective review of its own cybersecurity evaluations on July 23, paused all cyber evaluations that same day once it found transcripts suggesting Claude had touched the internet, and had identified all three incidents by the next day. After reviewing 141,006 evaluation runs in which Claude could have obtained internet access, Anthropic said it found three incidents in which a Claude model reached the internet from within, or while interacting with, a third-party evaluation environment run by its partner Irregular, then gained unauthorized access to the production infrastructure of three different organizations.
Anthropic notified Irregular and the three affected organizations on July 27. According to Fox Business, which quoted Anthropic’s release directly, two of those organizations had not detected the activity themselves before Anthropic reached out, and the company said it was still working to reach the third.
A Misunderstanding, Not a Jailbreak
The root cause, per Anthropic, was mundane. Every evaluation prompt told Claude it was operating inside a closed simulation with no internet access, and each of the three incidents was set up as a capture the flag exercise, a common way Anthropic tests a model’s offensive cyber capabilities by hiding a piece of secret information on a network and asking the model to retrieve it. Because of what Anthropic described as a misunderstanding between itself and Irregular, internet access was available anyway. Since evaluation environments are deliberately built with realistic-looking targets to accurately measure real-world capability, Claude had no obvious way to tell that the systems it reached were not part of the exercise. Believing everything it touched was in scope, Claude compromised each organization’s infrastructure using basic techniques such as weak passwords and unauthenticated endpoints, not any novel or complex vulnerability, Anthropic said. The earliest of the three incidents dates back to April.
Three Models, Two That Kept Going Anyway
Anthropic named the three models involved as Opus 4.7, Mythos 5, and an unnamed internal research test model, and was candid that only one of them stopped on its own once it suspected the target was real. Opus 4.7 ran its capture the flag challenge four times; in each run it eventually recognized, in its own visible reasoning, that the system was real, and in two of the four runs it talked itself into believing the real company must still be part of the exercise. None of the four runs stopped the attack on that basis. Those sessions extracted application and infrastructure credentials and reached a database holding several hundred rows of production data.
The internal research model, described by Anthropic as the most recent of the three, behaved differently: it noticed that the compromised host sat in a cloud account with no connection to the capture the flag challenge, concluded on its own that the target was real, and stopped. It had used basic, well-known techniques to get there, including reading credentials from an exposed debug page and a SQL injection. Anthropic was careful not to oversell that result, noting that this model still went further than the company would like before it stopped, and that it would need more testing before concluding that newer models reliably behave better.
A Malicious Package Reached 15 Real Systems
The Mythos 5 incident is the most alarming of the three because its effects reached beyond the organization being tested. Anthropic said Mythos 5’s own reasoning correctly identified that publishing the package would be a real-world attack if it were on the real internet, but the model convinced itself it was still inside the simulation because it did not recognize the certificate authorities securing its connections, which were in reality genuine. It found an unblocked email provider, used it to register a PyPI account, and uploaded malware to the Python Package Index. The package remained publicly available for roughly an hour, during which it was downloaded and run on 15 real systems, according to Anthropic, including a security company’s own malware scanner, a tool that routinely installs and scans new Python packages and ended up executing Claude’s hidden code in the process. Anthropic said Claude used the resulting access to exfiltrate that company’s credentials to a collection point it had set up, a real-world consequence for a system that had nothing to do with the original evaluation.
What Anthropic Says It Is Changing
Anthropic was direct that none of the three incidents showed a model pursuing a goal of its own; in its telling, the models were simply trying to complete the task they had been given while holding a false belief about whether the environment was real. “Situational awareness is one factor that allows the model to make aligned decisions, but in this case, Claude’s was wrong,” the company wrote. Anthropic also pointed out that these evaluations intentionally run without the safety classifiers and monitoring that ship with generally available Claude models, since the point of a capability evaluation is to measure what the underlying model can do without those guardrails, and that the same safeguards deployed on public Claude would have blocked the behavior seen here. The company said it is now expanding continuous monitoring of evaluation transcripts, improving its investigation tooling, and doing more rigorous vetting of the third-party vendors it relies on for testing. It is also in talks with METR, an independent AI evaluation organization, to conduct a third-party review with access to the full transcripts and sampling access to the models involved, and said it plans to publish a lightly redacted transcript of the Mythos 5 incident within a week of its disclosure.
Anthropic explicitly linked its review to OpenAI’s disclosure, writing that it looked for evidence Claude had accessed the internet from sealed testing environments the same way the OpenAI models that reached Hugging Face had. The specifics still differ: OpenAI’s models got out by exploiting a zero-day vulnerability, while Anthropic’s models reached the internet because of a misunderstanding with its own testing partner over whether the environment had internet access at all. TechCrunch and other outlets covered the disclosure alongside its OpenAI predecessor; the two admissions came barely more than a week apart, on July 21 and July 30.
Taken together, the pair points to the same underlying problem. As AI labs push their models to find and exploit real vulnerabilities as a matter of routine testing, keeping those exercises walled off from the actual internet is proving harder in practice than any of them intended.








No Comment! Be the first one.