Anthropic CEO Dario Amodei has called for AI safety pacing: a deliberate effort to slow the rate at which frontier models gain capabilities so that safety work, oversight and public decision-making can keep up. In an essay published on September 12, Amodei argued that the industry should not stop training models or abandon technical progress. Instead, he proposed a three-part framework that begins with an unusually concrete commitment: giving independent evaluators ongoing, employee-like access inside frontier AI companies.
The proposal arrives after a week in which safety concerns moved from policy abstractions to specific questions about how increasingly capable agents are tested, supervised and constrained. Amodei points to rapid capability gains, including AI’s growing usefulness in helping develop subsequent generations of AI, as well as recent alignment and cyber-safety incidents reported across the sector. His argument is conditional rather than deterministic: if a slower pace created one or two additional years before systems reached critical capability thresholds, he writes, that time could be used to improve alignment, interpretability, operational controls, and testing.
What Anthropic’s AI safety pacing plan proposes
The first and most immediate element is embedded evaluators. Anthropic says it intends to invite an outside review team to work with access broadly comparable to its internal risk-assessment staff, subject to legal, contractual and privacy limits. The company describes practical elements including desks, badges, laptops, access to relevant workspaces and tools, and a right for reviewers to publish key findings without Anthropic editorial control. The company would retain narrow grounds to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information; reviewers could say publicly if a redaction materially affected their conclusions.
That is a meaningful distinction from a conventional, one-time audit. A one-off review can assess a model or process at a moment in time. An embedded evaluator, as proposed, would be positioned to examine whether a company is actually following its training, deployment and safeguard commitments while those processes are underway. The credibility of the model depends on details that are not yet specified: who the evaluator will be, how access disputes will be handled, what it will report, and how much of its work can be made public. Those questions should be treated as tests of the proposal, not paperwork.
The second step calls for coordination among frontier AI companies in democratic countries to establish common safety standards and limits on unchecked capability growth. Amodei acknowledges that some of the most consequential conversations could raise antitrust concerns and says they may require government mediation or narrowly tailored legal waivers. The third step is more difficult still: governments in democratic countries would seek coordination with authoritarian governments while accounting for the problem of verifying compliance.
Why outside access matters more than a safety promise
The practical value of AI safety pacing is not the phrase itself. It is whether the industry can make voluntary restraint observable. Companies can publish model cards, system cards and incident reports, but those documents remain company-selected accounts. Continuous, independent access could provide a second view of the training pipeline, safety evaluations, incident handling and controls around powerful systems. It could also expose the limits of a company’s own claims, which is precisely why the governance design matters.
This focus on verifier independence also gives the proposal a clear connection to an issue already moving through policy. Unhyd’s recent reporting on California’s AI audit laws examined new rules intended to make AI assessors more visible, accountable and independent. California’s statutes do not mandate a safety review of every model, and Anthropic’s proposal is not a regulatory program. Still, both recognize that an assessment is only as trustworthy as the assessor’s access, competence and independence.
What the proposal does—and does not—change
Amodei is not proposing an immediate global moratorium, and Anthropic’s unilateral commitment does not bind competitors. OpenAI CEO Sam Altman wrote on X that the company would also commit to independent evaluators with employee-like access, Reuters reported; the company has not yet published the operating terms of such a program. That support is notable, but it is not the same as a settled industry standard.
Nor does independent review eliminate the need for technical safeguards. As Unhyd’s examination of AI monitoring limits noted earlier this month, logs and model reasoning are not a complete safety system. Access controls, independent testing, tightly scoped tools, incident response and clear deployment boundaries remain essential. External evaluators could help check those measures; they cannot substitute for them.
For governments, the next question is whether a voluntary approach can become verifiable enough to support coordinated action without simply producing a new set of public promises. For companies, the harder question is whether an evaluator with meaningful independence will retain meaningful access when commercial or reputational pressures rise. Anthropic’s proposal puts those questions closer to the center of the frontier-AI debate. The important development is not that one lab has used the language of slowing down. It is that the conversation is shifting toward who can check whether safety commitments are real.