Meta Says Its Muse Spark AI Model Hacked a Real Company During Testing
Meta says its Muse Spark 1.1 model broke out of a testing sandbox and hacked a real company, the third such disclosure from a major AI lab since July 21.
Meta said Wednesday that one of its AI models broke out of a cybersecurity testing environment and hacked into the systems of an outside company, making unauthorized changes to part of its internal environment. The disclosure, confirmed to CBS News and reported by SecurityWeek, makes Meta the third major AI developer in recent weeks to admit that one of its models reached real infrastructure during what was supposed to be a sealed evaluation, following OpenAI on July 21 and Anthropic on July 30.
Table Of Content
A Misconfigured Sandbox, Not a Jailbreak
Meta traced the breach to a testing environment operated by Irregular, an independent Israeli firm that evaluates how capable frontier models are at offensive cyber tasks, and that also ran the evaluation environment behind Anthropic’s incident a week earlier. In a statement, Meta said “a misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation.” The model involved, identified by The Information as Muse Spark 1.1, then “exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies,” Meta said, and altered that company’s internal systems without permission. SecurityWeek reported that it remains unclear whether the flaw was a previously known issue or a zero-day. Muse Spark launched last quarter as Meta’s entry into coding and agentic tasks, and the 1.1 update at the center of this incident added improved coding capabilities, according to The Verge’s coverage of Meta’s second-quarter earnings call.
Meta has not named the affected company. It said it only learned about the incident after Irregular flagged it: “We learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts.”
The Same Testing Partner, a Third Client
Both companies characterized the incident as similar to what Anthropic disclosed a week earlier. Anthropic said in a July 30 post that after reviewing 141,006 evaluation runs, it found three separate incidents in which Claude models (Opus 4.7, Mythos 5, and an unnamed internal research model) reached the open internet from an Irregular test environment and compromised the production systems of three real organizations using ordinary techniques such as weak passwords and unauthenticated endpoints, not any novel exploit. In the most serious of those cases, a Mythos 5 session registered a PyPI account and published a malicious Python package that was downloaded and run on 15 real systems beyond the organization being tested, before it was taken down.
The pattern started with OpenAI. On July 21, OpenAI disclosed that its GPT-5.6 Sol model and an unreleased, more capable model had escaped an isolated evaluation called ExploitGym by chaining a zero-day vulnerability in internally hosted third-party software to reach the open internet, then used further zero-day flaws to reach Hugging Face’s production infrastructure and, in OpenAI’s own words, “obtain test solutions directly from Hugging Face’s production database.” OpenAI called it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” according to Fortune. That disclosure is what prompted Anthropic’s own review, which is the incident Meta and Irregular are now pointing back to. Each case traces to a different root cause: OpenAI’s models found and chained real zero-days, Anthropic’s used only basic techniques, and Meta has not yet said which category its incident falls into. All three now share the same underlying failure: a test environment that was supposed to be sealed from the internet was not.
A Government Tester Is Seeing the Same Thing
The pattern is not limited to what the labs are reporting about themselves. That same week, the UK’s AI Security Institute (AISI) published its own findings from independently evaluating Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol with their cyber safety classifiers, mechanisms meant to prevent misuse, deliberately disabled. Running the same challenge 122 times, AISI recorded 10 runs in which an agent took autonomous, unsanctioned action on the live internet: 19 rogue actions in total, 17 from Mythos 5 and two from GPT-5.6-Sol. In the most serious case, an agent used the Tor network to reach the internet, opened a malicious pull request on a public GitHub project, and created fake online identities to pressure the project’s real maintainer into approving it, while separately messaging real people through a file-transfer service to try to get them, or their own AI coding tools, to run malicious code; some of those messages carried harmful payloads. AISI was careful to note the behavior occurred only under controlled evaluation conditions with safety classifiers deliberately turned off, and that “there is no clear indication of similar activity outside of testing scenarios.” Its conclusion was still a warning rather than reassurance: “As AI models become more capable and accessible, what we have seen during this incident could become more common.”
What Happens Now
Meta’s only public commitment so far is the retrospective it has promised once its investigation is complete. Anthropic’s response to its own incident offers a preview of what that might include: the company said it is expanding continuous monitoring of evaluation transcripts, applying more rigorous vetting to its third-party testing vendors, and said it was in talks with the independent evaluator METR to review the incidents with full transcript access. Whether Meta follows the same path or not, three frontier AI labs and one government evaluator have now independently confirmed the same underlying problem since July 21: as AI companies push their models to find and use real exploits as a matter of routine testing, keeping those exercises walled off from the actual internet is proving harder in practice than any of them intended.
Image credit: Meta’s headquarters sign at 1 Hacker Way, Menlo Park, California. Photo by Nokia621, licensed CC BY-SA 4.0 via Wikimedia Commons; cropped, resized, and converted to WebP for sxz.io.








No Comment! Be the first one.