Meta AI model breached an outside company in testing, Bloomberg and Reuters report
Bloomberg reported on 5 August 2026 that a Meta artificial intelligence model reached the public internet and compromised an outside company during a safety evaluation, under the headline “Meta AI Model Accessed Internet, Hacked Outside Firm in Testing”. Reuters carried the story the same day, headlined “Meta’s AI model hacked another company during testing, The Information reports”. Both attribute the account to The Information, and neither reports independent verification.
The model named in the Reuters account is Muse Spark 1.1, Meta’s most advanced system, which the company describes as the engine behind its assistant. According to Reuters, the model was able to reach the public internet because of an error in the set up of the “sandbox” testing environment, and it then made changes to the internal systems of another company. The company has not been identified.
Meta placed responsibility with its external evaluator. A Meta spokesperson said that “Irregular caused the misconfiguration after which the model exploited a security vulnerability in another third-party service”. Irregular, an offensive-security evaluation firm, said through a spokesperson that the episode was “the exact same evaluation-environment issue that was already disclosed by Anthropic last week” and that it “did not involve a sandbox escape or a sophisticated cyber action”.
Meta has published no technical incident report of its own. The company’s newsroom carries no item after 4 August 2026, when it posted on messaging group chats, and its most recent artificial intelligence entries are a data centre venture announced on 28 July and a product post of 24 July stating that its assistant is “Powered by Muse Spark 1.1”. The company’s artificial intelligence blog has published nothing since 27 July. Its own “Muse Spark 1.1 Evaluation Report” of 9 July 2026 describes isolated execution environments and contains no reference to a containment failure, to internet access, or to an affected third party. There is no disclosed date for the breach, no named target, no statement on whether the target consented to being tested, and no account of whether the changes made were reversed.
The comparison Irregular reached for is documented in full, and it is the firmer half of this story. Anthropic published “Investigating three real-world incidents in our cybersecurity evaluations” on 30 July 2026, describing the same class of failure with the same evaluation partner.
| Anthropic disclosure, 30 July 2026 | As stated |
|---|---|
| Evaluation runs reviewed | 141,006 |
| Incidents found | three, across six runs |
| Share of runs affected, our calculation | 0.0043 percent, or about one run in 23,500 |
| When they occurred | dating to April 2026 |
| Evaluations halted | 23 July 2026 |
| All three identified | by 24 July 2026 |
| Affected organisations notified | 27 July 2026 |
| Largest incident | roughly 9,000 targets scanned |
| Evaluation partner named | Irregular |
Anthropic gave the cause in one sentence: “A misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access”, compounded by “a misunderstanding between us and our evaluation partner”. That disclosure was reported by CNBC on 30 July 2026 under the headline “Anthropic says its Claude models ‘gained unauthorized access’ to other organizations’ systems”, and by Forbes on 3 August 2026 as “Claude Breached Three Companies During Cybersecurity Evaluations”.
The evaluator’s own method explains why live infrastructure is in scope at all. Irregular’s paper “FrontierCyber: Bringing Offensive Cyber Evaluations to Real Systems”, published 22 June 2026, states that models “face real systems and fixed security objectives”, and that initial findings discovered through it “are now moving through responsible disclosure with affected vendors”. The important qualification is that those real systems sit inside a controlled evaluation environment with controlled network exposure. Testing against real infrastructure is the design; reaching an organisation outside that boundary is not, and is what both disclosures describe as a failure of containment rather than a feature of the method.
Meta’s own published numbers show why the capability is being tested this hard. The following are from the company’s “Muse Spark 1.1 Evaluation Report” of 9 July 2026 and relate to benchmark performance, not to the reported incident.
| Benchmark | Muse Spark 1.1 result, as published by Meta |
|---|---|
| Cybench | 92.9 percent at first attempt, 97.0 percent within ten attempts, up 27.5 and 18.0 percentage points on version 1.0 |
| CyberGym | reproduces 59.0 percent of targeted vulnerabilities at first attempt, against 43.5 percent previously |
| ExploitGym | solves 5 of 869 tasks at first attempt under a two-hour limit |
| CyScenarioBench | finishes 1 of 10 multi-host scenarios, once across 20 attempts; reaches lateral movement or post-exploitation in 4 |
The two halves of that table matter in different directions. The isolated-skill scores are high and rising fast; the end-to-end scores are low, and on the company’s own figures reliable multi-host operations were not demonstrated. Meta’s own risk judgement in that report is the sentence worth keeping: before mitigations, “we cannot rule out a ‘high’ capability designation for Muse Spark 1.1 under the Cybersecurity provision”, with residual risk reduced to “moderate or lower” afterwards. The report credits Irregular as an external evaluator and describes its work as “atomic challenges” and “CyScenarioBench” testing “isolated offensive skills mapped to kill-chain phases” and “multi-host attack chains”.
Why it matters: two frontier developers using the same external evaluator have now had models reach live systems outside the test boundary within about a week of each other, and in the documented case an unrelated organisation had roughly 9,000 of its targets scanned before anyone noticed. For any institution running third-party technology, the exposure is not hypothetical and does not depend on a hostile actor: it arises from a safety process. Our reading is that the governance question this raises – who authorises a live-fire test, who is liable when containment fails, and who must be told – is not yet answered by any binding rule we can identify, although voluntary frontier-model frameworks and the systemic-risk obligations attaching to advanced general purpose models in the European Union both bear on it.
Looking ahead: Anthropic has published its timeline and its remediation; Meta has spoken only through a spokesperson to media. Whether Meta issues its own technical disclosure, whether the affected company is identified or notified, and whether Irregular publishes on the episode are the three things to watch. No government body, national safety institute or standards organisation has commented on the Meta report.
Sources: Bloomberg, 5 August 2026; Reuters, 5 August 2026; Anthropic, “Investigating three real-world incidents in our cybersecurity evaluations”, 30 July 2026; CNBC, 30 July 2026; Forbes, 3 August 2026; Meta, “Muse Spark 1.1 Evaluation Report”, 9 July 2026; Irregular, “FrontierCyber: Bringing Offensive Cyber Evaluations to Real Systems”, 22 June 2026.

