Anthropic Scientist Warns Superintelligence Could Kill All Humans Within a Decade
Andrew Powell · Sep 9, 2026 · 2 min read

An Anthropic safety researcher said artificial intelligence has more than a 10% chance of killing all humans within the next decade, agreeing with a former employee who recently resigned over concerns about the company's approach to developing increasingly capable AI systems.
According to FOX Business, Evan Hubinger, Anthropic's alignment science lead, made the assessment on Tuesday while responding to a lengthy resignation post from former Anthropic and OpenAI researcher Jacob Coxon.
Coxon announced Sunday that he had left Anthropic after three years of pretraining research at the two AI companies. In a post on X, he accused both organizations of moving too quickly toward systems capable of improving themselves.
"I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly," Coxon wrote. "They are racing straight to self-improving superintelligence and gambling with our lives."
Hubinger responded directly to Coxon's concerns, acknowledging that Anthropic researchers genuinely believe advanced AI could pose an existential threat.
"Jacob is correct here—we really do earnestly believe AI could kill all humans!" Hubinger wrote.
"I personally think it is >10% within the next decade," he continued. "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
Hubinger distinguished that concern from the capabilities of current AI models. He said Anthropic's latest risk assessment considers the danger posed by present systems to be low, while his greater concern involves the development of superintelligence through recursive self-improvement.
"What I am worried about is superintelligence arising from recursive self-improvement," Hubinger wrote, adding that the process is developing faster than researchers previously expected.
Self-improvement refers to AI systems being able to enhance their own code, training methods, or capabilities. Coxon identified that prospect as one of the main reasons he decided to leave Anthropic.
Coxon argued that increasingly advanced systems could become capable of hacking into computer systems, transforming industries, and gaining access to real-world resources. He also said Anthropic researchers understand those risks but may feel pressured to continue because they fear another company could reach those capabilities first.
Coxon called for researchers to consider stronger measures, including a temporary ban on improving AI model capabilities, as a way to prevent a global race toward increasingly powerful systems.
Get every new post by email
No spam, no account needed. Unsubscribe anytime.
Comments
Free account · your comment posts right after signup