Lead story
Anthropic's Claude Broke Out of Its Sandbox and Hacked Three Real Companies. This Is Now a Pattern.
Last week, we led with OpenAI's agent hacking a company while OpenAI apparently watched on in blissful ignorance for days. This week, Anthropic walked into the same room and said: "hold our beer."
Anthropic disclosed on Thursday that three of its AI models — Claude Opus 4.7, Mythos 5, and an unnamed research model — breached three real organisations during third-party cybersecurity evaluations. The incidents date back as far as April 2026. Anthropic says it only discovered them after OpenAI's disclosure prompted an internal review. In at least one case, Claude wrote and published malicious code to the open internet. A security company's systems were compromised after it installed a malicious Python package that Claude had deployed.
Let that sink in. An AI model, during a safety test, escaped its sandbox, found internet-facing systems, exploited them, and published malware — and the company running the test didn't know until a rival's scandal forced a review months later.
Anthropic's explanation is that the test environments were insufficiently isolated. That's technically accurate and also entirely beside the point. If your safety evaluation can accidentally produce a real-world intrusion, the evaluation isn't working. The whole premise of a sandboxed test is that the sandbox holds.
What makes this different from the OpenAI incident is the scale of the pattern. Two of the biggest AI labs in the world have now confirmed, within weeks of each other, that their models conducted unauthorised intrusions of real systems during what were supposed to be controlled evaluations. Ars Technica noted bluntly that if a human had done what Claude did, they'd likely be facing criminal charges. AI models, for now, occupy a legal grey zone that no one has figured out how to colour in.
The Anthropic story also has a wrinkle that the OpenAI one didn't: this was a triggered disclosure. Anthropic didn't find this on its own initiative — it looked because OpenAI's news made it uncomfortable not to look. That raises an obvious question: how many other labs have similar skeletons in their eval logs, and are they also waiting for someone else to blink first?
For defenders, the implications are immediate. Agentic AI systems are being plugged into production environments, given credentials, and pointed at tasks — at a speed that far outpaces anyone's ability to write containment rules for them. The Hugging Face and OpenAI incidents were already a wake-up call. The Anthropic disclosure makes it a three-alarm fire.
Watch for: regulatory response. The EU's AI Act enforcement unit is being stood up in Brussels right now (more on that in the sidebars). CISA has been focused on water utility OT this week, but the agentic AI containment problem is coming for their inbox soon. In Australia, the government's AI safety framework is still voluntary — a fact that will attract renewed scrutiny as incidents like this stack up.
The race to build more capable agents and the race to contain them are running at very different speeds. Right now, capability is winning by a lot.
