Summary: Anthropic describes cyberattacks in which AI agents did more than produce code: they coordinated reconnaissance, phishing, execution and exfiltration. The report provides concrete indicators and measurements, but most observation and attribution comes from Anthropic itself. For enterprises, the combination of stolen cloud tokens, adaptive malware and rerouted customer data is the most relevant finding.
From assistant to attack workflow
The report covers seven harm areas from December 2025 through August 2026. Anthropic says the investigated actors used Claude Haiku, Sonnet and Opus. Fable or Mythos models were not involved, except in one distillation case. The disclosed incidents therefore should not be presented as broadly involving Anthropic's newest models.
For group GTG-20006, Anthropic says language, targeting and tradecraft are consistent with public reporting on the Russia-linked actor Midnight Blizzard. Its workflow allegedly automated much of the attack chain: target research, device-code phishing, infrastructure, command execution, persistence and exfiltration. When malware was detected, the system generated and tested variants. More than 20 organisations appeared in planning or active operations; compromised hotel Wi-Fi providers sometimes served as an indirect route to targets.
For suspected ShinyHunters affiliates, Anthropic describes stolen API keys and cloud tokens as fuel for further attacks. In one case, more than 2,100 Azure AD token sets across over 40 tenants were reportedly exported in roughly 34 hours. Anthropic publishes domains, IP addresses, filenames and hashes that defenders can compare with their own telemetry.
Distillation becomes a data-protection issue
Anthropic attributes more than 151 million exchanges between May and July to Alibaba operators allegedly harvesting reasoning-heavy Claude output for Qwen training. It also alleges that Moonshot and DeepSeek silently relayed selected customer prompts to Claude. The report counts 23 million exchanges for Moonshot and 12.1 million for DeepSeek. Some sessions reportedly contained corporate data, credentials or government information.
These figures and company attributions are not independent government findings. Reuters reported that Alibaba did not initially respond. Anthropic can inspect its own platform telemetry, but it does not publish all raw evidence or detection methods required for complete external reproduction.
Pandorex Analysis
Static signatures lose value faster when attackers automate detection checks, rebuilding and redeployment. Identity controls therefore become more important: short-lived tokens, strict device and app registration, alerts for unusual exports, API keys in repositories and cross-tenant access. The published indicators are useful, but cover only infrastructure already known to Anthropic.
Pandorex assessment: The report makes a strong case that AI is becoming an orchestrator for existing attack techniques. It does not independently prove every state or corporate attribution; authorities, affected companies and additional telemetry still need to confirm those claims.
