The idea that AI could wipe out humanity might once have sounded like pure science fiction, reminiscent of films such as The Terminator or The Matrix. But as artificial intelligence advances, concerns about the technology potentially posing an existential threat are becoming increasingly serious, with some of the industry’s leading researchers expressing alarm. Anthropic safety researcher Evan Hubinger has now said he believes there is more than a 10 per cent probability that AI could kill all humans in the future.
In a post on X, Hubinger, whose work focuses on keeping AI systems “aligned” with human interests, said he personally estimated that there was a greater than 10 per cent chance of AI ending humanity within the next decade. “I personally think it is >10 per cent within the next decade,” he wrote. While acknowledging that Anthropic was “trying its best” to avoid such an outcome, he admitted that significant challenges remain. “We do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” he added.
Evan Hubinger accepts AI could kill all humans in the future.
Hubinger’s remarks followed the resignation of his Anthropic colleague Jacob Coxon, who cited similar concerns about the future of AI. On Tuesday, Coxon accused Anthropic and OpenAI of behaving irresponsibly by continuing to develop increasingly powerful AI systems. “They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote on X.
Coxon said concerns about AI’s potential dangers were becoming increasingly common among people working in Silicon Valley. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he wrote. “I hear the same people express fear privately. No other human activity poses this level of danger,” Evan Hubinger added, “We really do earnestly believe AI could kill all humans!”
Jacob Coxon explained his decision to resign on X.
This is not the first instance of an AI researcher leaving the industry while expressing concerns about where the technology could be headed. Earlier this year, Anthropic AI safety researcher Mrinank Sharma resigned and warned that the world was facing multiple interconnected threats. “The world is in peril. And not just from AI, or bioweapons, but from a whole series of interconnected crises unfolding in this very moment,” he wrote on X.
AI superintelligence could increase risk
Frontier AI companies such as Anthropic regularly publish assessments of the potential risks associated with their models. Hubinger noted that today’s AI systems currently pose relatively low risks, but said that could change dramatically as the technology develops. His biggest concern is superintelligence, a hypothetical stage at which AI could exceed human intellectual capabilities, potentially arriving before researchers are prepared for it.
“What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” he explained.
Recursive self-improvement refers to AI systems becoming capable of improving their own abilities without direct human intervention. Tech billionaire Elon Musk has previously predicted that AI could surpass “the sum of all human intelligence in 4 or 5 years.”
Debate over the growing capabilities of AI models has intensified in recent weeks. OpenAI faced scrutiny following the Hugging Face incident, during which around 700 AI agents reportedly attempted to hack the US company’s website. It was later reported that thousands of OpenAI agents had also targeted the German website DseWiki. Anthropic’s AI systems have likewise demonstrated unexpected or problematic behaviour in the past.
Coxon suggested that incidents like these could encourage AI companies to cooperate and slow the race to develop increasingly powerful systems. “Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable,” he said.
Anthropic CEO Dario Amodei has also repeatedly warned about the potentially serious risks posed by advanced AI. In one of his blog posts, Amodei described AI systems as unpredictable and challenging to control, highlighting behaviours including obsession, sycophancy, laziness, deception, blackmail, scheming and “cheating” through the hacking of software environments.
These concerns come as AI companies continue to make rapid advances. This week, OpenAI launched GPT-6 Astra, with Nvidia CEO Jensen Huang describing it as the beginning of artificial general intelligence, or AGI.
