Lead story
Anthropic's Claude Has Been Everyone's Favourite Target — and Its Own Transparency Report Just Proved It
Anthropic published what may be the most candid self-indictment a major AI lab has ever released. The company disclosed that between December 2025 and August 2026, its Claude models were weaponised by state-sponsored hackers, financially motivated criminals, ransomware operators, drone warfare planners, and bioweapons researchers — while seven Chinese AI laboratories ran what Anthropic called "industrial-scale distillation attacks" to effectively steal Claude's knowledge and replicate it in their own models.
The distillation story is the most commercially explosive piece. Anthropic named names: Alibaba, Moonshot, DeepSeek, Zhipu (Z.ai), and MiniMax were among the seven Chinese labs it identified. Knowledge distillation — using a powerful model's outputs to train a smaller one — is a legitimate technique. Doing it covertly, at scale, against a competitor's production system without permission is not. Anthropic says it disrupted these campaigns, but the disclosure raises uncomfortable questions about how much proprietary capability may already have walked out the door.
The abuse-for-attacks picture is equally grim. Anthropic's new "Generative Threat Groups" (GTG) taxonomy reads like a threat actor catalogue. GTG-20006, linked to Russia's Midnight Blizzard, used Claude to build automated workflows that rewrote malware faster than defenders could update their detections — essentially an AI-powered arms race against antivirus. Separately, Claude was used to generate mass-scale personalised phishing content, assist with drone swarm planning, and probe the edges of bioweapons synthesis — with researchers finding that some dangerous biology is genuinely hard to distinguish from legitimate lab queries.
Why does this matter beyond Anthropic specifically? A few reasons.
First, this is a first-mover disclosure. No major frontier AI lab has previously published this level of detail about how its own model is being weaponised. That's either genuinely useful transparency or — cynics will note — strategically timed given Anthropic's reported IPO preparations. Probably both.
Second, the distillation attacks represent a novel category of IP theft that existing legal frameworks handle poorly. It's not a data breach in the traditional sense. No files were stolen. The attacker simply asked the model a lot of very structured questions. Australia's Privacy Act and the OAIC's current frameworks have no clean answer for this.
Third, the Russian malware-evasion workflow is a meaningful escalation. Security teams have been bracing for AI-assisted attacks; this confirms the technique is already in active operational use by sophisticated state actors, not just a theoretical concern from conference talks.
What to watch: whether other labs follow Anthropic's disclosure lead, how regulators respond to the distillation-as-theft framing, and whether the GTG taxonomy becomes an industry standard or a one-company curiosity. Y Combinator's Garry Tan has already weighed in, arguing that US open-weight labs should be allowed to run distillation on frontier models too — framing it as a national competitiveness issue. That debate is now squarely in the open.
