Anthropic Researcher Says AI Has More Than 10% Chance of Killing All Humans Within a Decade
Key Takeaways
- •Evan Hubinger said he estimates the chance of AI killing all humans within the next decade at more than 10%.
- •Hubinger said the immediate danger is not from current models but from superintelligence developed through recursive self-improvement.
- •Hubinger acknowledged that Anthropic does not yet have a plan to solve superintelligence alignment or a clear path to doing so.
- •Former researcher Jacob Coxon resigned from Anthropic after accusing it and OpenAI of recklessly pursuing self-improving AI systems.
- •Coxon said a temporary ban on improving model capabilities may be needed to prevent a dangerous global race.

A senior Anthropic safety researcher said Tuesday that artificial intelligence has a greater than 10% chance of “kill[ing] all humans” within “the next decade,” responding to a former employee who resigned after accusing the company of acting irresponsibly.
Former Anthropic and OpenAI researcher Jacob Coxon wrote in a lengthy resignation thread posted to X on Sunday that “the people building AI earnestly believe that it could kill us all by the end of the decade.”
“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” Coxon wrote.
Evan Hubinger, Anthropic’s alignment science lead, responded to Coxon’s thread in a quoted post. Hubinger said Coxon’s assessment was correct, while adding caveats about the source and timing of the risk.
“Jacob is correct here—we really do earnestly believe AI could kill all humans!” Hubinger wrote. “I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Hubinger said the potentially deadly risk does not come from current AI models. “To be clear, as we say in our latest Risk Report, I think the risk from present models is low,” he wrote. “What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.”
Self-improvement refers to an AI model’s ability to continuously enhance its own source code or training methods. Coxon specifically cited research into this capability as a primary reason for his resignation.
“These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing,” Coxon wrote in his X thread.
Coxon said Anthropic’s scientists understand the risks of their work but continue advancing because they fear a less responsible company could unlock those potentially disastrous capabilities first.
“At Anthropic, the stakes are well-understood, but they are locked in a race to get there first—they believe no one else will act responsibly, so they must do it themselves, despite the risk,” Coxon added.
To prevent what he described as an absolute catastrophe, Coxon suggested that the world might need a temporary ban on improving model capabilities.
“I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities,” he wrote.
Coxon ultimately urged AI researchers and developers worldwide to consider the consequences of their work and act more responsibly.
“If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’—or take this moment to call for different conditions?” his X thread concluded.
FOX Business reached out to Coxon, Hubinger, Anthropic and OpenAI for further comment.