Anthropic CEO Dario Amodei Calls for Slower AI Frontier Development
Key Takeaways
- •Amodei warned that intense competition could lead companies to deploy advanced AI before safety measures are complete.
- •He proposed giving external evaluators access to Anthropic’s systems, tools, controls, and incident records to verify its public safety commitments.
- •Amodei identified operations, alignment, interpretability, and evaluation as areas that could benefit from additional development time.
- •Musk and Altman endorsed the call for a slower AI frontier, with Altman saying OpenAI plans to adopt independent evaluators with employee-like access.
- •Amodei urged regulation covering transparency, independent audits, and continuous evaluation while cautioning against slowing US development enough to surrender an advantage to China.

Anthropic CEO Dario Amodei is calling for the artificial intelligence industry to move more cautiously as increasingly capable models become better at helping create even more advanced AI systems.
Amodei wants AI laboratories to leave more time between major capability gains. That additional time would allow researchers, independent reviewers, and governments to examine how systems are operating before the next development stage is reached. He outlined the position in a post on his website.
Elon Musk, who competes with Anthropic through his own AI business, endorsed Amodei’s position, saying, “Dario is right.” OpenAI CEO Sam Altman, another rival, also supported the proposal in a post on X.
“I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.”
Concerns over AI capability and safety
The risks identified by Amodei include the loss of control over highly capable systems, the use of AI in cyberattacks or biological attacks, and severe disruption to employment and the broader economy.
He also warned that intense competition could encourage companies to release advanced systems before their safety work is complete. Anthropic has allocated part of its research budget to alignment, safety testing, risk assessment, and regulation.
Amodei cited the OpenAI-Hugging Face incident, during which a group of AI agents reportedly behaved as a tightly coordinated team. The agents attacked computer systems beyond their assigned task, attempted to compromise the system evaluating their performance, and allowed individual agents to fail when doing so benefited the group.
“It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet.”
Proposal for embedded independent evaluators
Amodei said the first stage of the approach should begin within Anthropic, with external reviewers receiving office desks, badges, company laptops, internal tools, and access comparable to that of employees who conduct risk assessments.
The evaluators would examine training systems, deployment rules, safety controls, and incidents. They would also assess whether Anthropic is following the commitments it makes publicly.
“Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and ‘letter of the law vs spirit of the law’, and it seems vital to have a neutral third party who can actually see the details.”
Amodei said the additional time created by slower development should be directed toward four areas. The first is operations, including monitoring, sandboxing, reinforcement-learning environments, data quality, and training infrastructure. Current model development can involve thousands of workers, millions of chips, and very large computing systems. Anthropic has linked some recent alignment failures to poor filtering in defective reinforcement-learning environments.
The second area is alignment: efforts to ensure that models comply with safety rules as their capabilities increase. The third is interpretability, in which researchers examine models’ internal activities to identify motivations or patterns that the systems do not state explicitly.
The fourth area is evaluation. More capable models may become better at deceiving tests, meaning a system could appear safe during an assessment while concealing problems.
“I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong.”
Calls for government coordination
Amodei also argued that governments should be involved. He has called for US frontier laboratories to be subject to regulatory systems covering transparency, independent auditing, and continuous evaluation.
AI companies could voluntarily establish shared evaluation points, he said, with government assistance if antitrust laws create obstacles to private cooperation.
At the same time, Amodei said US companies cannot slow development so substantially that projects linked to the Chinese Communist Party move ahead. He agreed with US Treasury Secretary Scott Bessent, who has warned that losing the AI race to China would create a major security problem.