In brief: Anthropic is combining Project Glasswing and its Cyber Verification Program into three vetted access tiers. Authorised red teams receive Claude models without the usual cyber blocks, while testing life-critical systems remains limited to a small group reviewed with US authorities. The disclosed figures show substantial reach, but not independently measured security impact.
Three boundaries instead of one global filter
Defense Access covers incident response, malware analysis and vulnerability validation. Organisations, infrastructure operators, open-source maintainers and researchers with a disclosure record may apply. Red Team Access adds authorised penetration testing for organisations, while still blocking physical harm or mass disruption.
Specialized Access has the fewest blocks and is reserved for deeply vetted organisations testing systems such as power grids, flight operations or interbank transfers. Anthropic says it reviews participants with the US government. Every tier includes Opus 5.5, Sonnet 5.5 and Mythos 5.1 on the Claude Platform, Google Vertex AI and Microsoft Foundry. Amazon Bedrock is initially limited to customers eligible for Enterprise Frontier Safeguards.
The evaluation measures blocks, not finding quality
Anthropic ran Opus 5.5 five times on ten multi-stage CyScenarioBench tasks. Without programme access, safeguards stopped every task at the first prompt. Defense Access blocked 46 of 50 trials and allowed four to succeed. Red Team Access triggered no blocks and completed 34, almost matching the stated 67.6% without safeguards.
This shows that the tiers behave differently. It does not establish how many findings are valid, how often legitimate work is stopped, or whether authorisation is enforced reliably outside the model. Identity checks, target authorisation, logging and downstream approval therefore matter at least as much as the model classifier.
Pandorex Analysis: impressive volume, unknown patch rate
Glasswing partners reported at least 129,000 verified vulnerabilities from April through July; Anthropic's open-source scans found another 5,500 by October. More than 33,000 were rated critical or high. The scale matters, but figures come from heterogeneous reports by 33 partners. Fewer than half disclosed patch counts, while the estimate of at least five times greater total impact remains an extrapolation.
The new access model complements commercial continuous red teams using several frontier models. Vendors are moving the security boundary from blanket refusal towards verified identity and tiered permissions. That may give defenders more capability, while shifting part of the risk into admission checks, monitoring and misuse response.
