Anthropic’s Own Alignment Lead Says There Is A 10% Chance AI Kills Everyone Within A Decade

Things got remarkably candid on X when Evan Hubinger, who heads alignment science at Anthropic, admitted he sees a more than 10% chance that AI wipes out humanity within ten years.

The comment came as a reply to resigning colleague Jacob Coxon, who warned that frontier labs are recklessly playing with fire. Rather than trying to soften the blow, Hubinger leaned right into the critique and gave it a timeline and a metric, specifically tying his estimate to the risk of recursive self-improvement.

The weight of the comment comes from Hubinger’s position at Anthropic. He leads the division tasked with ensuring advanced AI remains safe and obedient as it evolves. And when the person in charge of alignment at an industry-leading safety lab acknowledges these odds, it acknowledges a glaring problem: the sector is racing towards superintelligence without a solution to keep it under control.

 

Where Does Recursive Self-Improvement Fit In?

 

Hubinger’s timeline, thankfully, doesn’t apply to the tools currently on the market, which he views as largely manageable. Instead, he’s looking at superintelligence born out of recursive self-improvement, something he claims is improving faster than the industry anticipated.

It all boils down to a feedback loop: AI systems get smart enough to help build their successors by refining algorithms, designing better hardware tools and optimising training methods. Each upgraded generation then takes over the task of engineering the next, drastically shortening development cycles. Left to run, that loop could pull off massive capability leaps, ultimately producing systems that outclass human intellect in almost every field.

Anthropic has previously noted in its own risk assessments that unconstrained self-improvement increases the odds of humans losing control over AI altogether. This official documentation connects the technical reality to the loss-of-control scenarios Hubinger is airing in public. Today’s chatbots are largely benign, but this automated cycle threatens to collapse safety timelines completely.

The risk of extinction is fundamentally about control, not malice. Advanced software accelerating its own development doesn’t need bad intentions to prove catastrophic. A system simply needs to reach a level of capability where people can no longer monitor its optimisation targets or redirect its trajectory if its objectives diverge from human interests.

 

What’s Driving the Sudden Urgency?

 

A combination of factors is driving these concerns. Progress on AI systems capable of aiding their own engineering is outpacing previous expectations, bringing a functional self-improvement cycle closer to reality. Simultaneously, Hubinger has made it clear that solving alignment for superintelligent systems is still an open challenge. The timing of his statement, paired with an engineer’s public exit, points to seemingly widespread unease among technical insiders.

Hubinger also later reinforced that current commercial models are safe, focusing his warnings on the trajectory towards self-improving systems. This nuance is key to understanding his assessment. An insider with direct line of sight into cutting-edge capabilities is pointing out that the field is speeding down a track without a clear brake mechanism.

 

What This Does To Public Trust

 

Direct warnings of unsolved, existential risk from an alignment lead make it exceptionally difficult to write off safety concerns as theoretical noise.

Anthropic actively trades on its reputation as a responsible developer. Seeing senior engineers at the heart of that work attach high probabilities to total disaster speaks volumes about the actual state of internal progress. It also exposes a growing contradiction for companies operating at the forefront. Anthropic continues to pitch the virtues of increasingly powerful systems while its own safety head assigns double-digit odds to human extinction within ten years.

Industry observers note that while Hubinger’s metric represents his individual assessment rather than formal corporate policy, it echoes sentiments held across Anthropic’s research divisions.

From a regulatory perspective, such explicit warnings from senior technical insiders certainly build a compelling argument for binding government oversight and mandatory safety checks before releasing self-improving models, and replacing voluntary self-policing with enforced accountability.