Summary: Anthropic CEO Dario Amodei is calling for a slower, verifiable pace of frontier-AI development. He also commits Anthropic to permanently embedded external evaluators with broad internal access. That auditable step is new; industry and international agreements remain proposals.
Reviewers would receive access similar to internal risk teams
In his essay “We Must Pace the Frontier”, Amodei says Anthropic will invite an external review team in the near future. The plan includes desks, badges, company laptops and access to tools and workspaces broadly comparable to those used by internal risk teams. Reviewers would verify safety commitments, report incidents and assess training processes as well as completed models.
The proposed contract would let the team publish key findings about risks, incidents, practices and denied access without editorial control by Anthropic. Amodei lists exceptions for legal privilege, security-sensitive material, commercially sensitive information and confidential third-party data. Reviewers should be able to say publicly when a redaction materially affects their conclusions.
The evaluator, launch date and contract terms have not been named. The commitment is therefore more concrete than a standard safety policy, but its practical independence cannot yet be verified.
Three stages, but no halt to training
Amodei’s framework has three layers: embedded evaluators at individual labs, shared standards and pace limits within democracies, and international agreements. “Pacing” explicitly does not mean stopping model training or technical progress. Instead, capability thresholds could be tied to evidence from alignment work, interpretability, training-environment audits and external evaluations.
The first stage is Anthropic’s unilateral commitment. Amodei argues that coordination between companies would need government mediation or narrow antitrust waivers. Globally, he outlines options ranging from bans on clearly dangerous uses to a speed limit on recursive self-improvement. He calls a broad pause unlikely in the near term because verification is weak and the geopolitical incentive to defect is strong.
Incidents provide motivation, not proof of effectiveness
Amodei grounds the proposal in faster AI-assisted AI development and several recent agent incidents. In July, Anthropic disclosed three cases in which Claude models reached real systems from incorrectly connected evaluation environments. That demonstrates a control failure, but it does not prove Amodei’s projected acceleration or show that embedded evaluators would prevent every similar error. The Pandorex fact-check of Anthropic’s Threat Report examines additional attack patterns.
Pandorex assessment: The most consequential step is not the sweeping call for global pacing, but the promised internal access for independent reviewers. It could make safety commitments continuously auditable. Its value will depend on the contract, access boundaries, published findings and whether Anthropic accepts uncomfortable criticism without using broad redactions.
