AI Existential Risk: Anthropic’s 10% Warning Explained

Anthropic AI existential risk warning and over 10% chance of catastrophic AI outcome
AshrafulIslam Avatar

Talk of AI existential risk just moved from message-board speculation to something Anthropic’s own safety staff are saying in public, under their real names, this week.

Why Anthropic’s Own Researchers Are Raising the Alarm. On Tuesday, Anthropic researcher Jacob Coxon announced on X that he was resigning, accusing leading AI labs of racing toward superintelligence without adequate safeguards. Hours later, Evan Hubinger, who leads alignment science at Anthropic, backed him up and put a number on it: more than a 10% chance of a catastrophic outcome within the next decade.

Quick answer: Two Anthropic employees, one departing and one still at the company, publicly stated this week that they believe advanced AI carries a real chance of causing mass human harm, with Hubinger estimating that chance at over 10% within ten years. Anthropic has not disputed either statement.

Step 1: Jacob Coxon Resigns and Sounds the Alarm

Coxon, who had worked on pretraining research at both OpenAI and Anthropic, posted his resignation announcement on Tuesday. He wrote that “racing straight to self-improving superintelligence and gambling with our lives” is exactly what the leading labs are doing right now.

He also told the Wall Street Journal that the most aggressive scenarios could see AI development slip out of human control as early as the end of 2027.

That’s a specific, testable claim, not a vague warning. Coxon isn’t some outside critic either. He’s someone who built pretraining systems at two of the companies now leading the race, which is part of why his resignation got picked up so fast by Axios, CNBC, and Forbes within hours.

Step 2: Evan Hubinger Confirms the 10% Estimate

What made this story move from “one researcher quits” to “Anthropic has an internal crisis of confidence” was Hubinger’s response. He didn’t distance the company from Coxon’s claims. Instead, he confirmed them directly: “We really do earnestly believe AI could kill all humans.”

READ MOREBest AI Tools for Beginners: 7 Easy Picks for 2026
Best AI Tools for Beginners: 7 Easy Picks for 2026

Hubinger added that he personally puts the odds above 10% in the next decade, and that Anthropic is “trying its best” but doesn’t yet have a working plan to solve alignment for superintelligent systems. He also said the company is “not clearly on track” to solve it. Coming from the person whose actual job is alignment science, that’s a notably blunt admission.

Step 3: What This Existential Risk Warning Actually Means

It’s worth being precise about what’s being claimed, because talk of AI wiping out humanity sounds like sci-fi shorthand until you unpack it. Both researchers are talking about a specific failure mode: AI systems capable of building their own successors, improving themselves faster than humans can monitor or correct them.

Anthropic itself addressed this scenario in a blog post, stating that once systems can fully build their own successors, the ways researchers secure, monitor, and shape their behavior become far more important.

Hubinger also clarified in a follow-up post that he’s not worried about today’s models causing this outcome. He described the risk posed by Anthropic’s currently deployed systems as low, and said his concern is specifically about superintelligence that could emerge from recursive self-improvement, not the chatbots people use right now.

Step 4: How This Connects to the July Hugging Face Incident

Neither researcher’s warning came out of nowhere. Reporting from CNBC and CNBC Africa notes that concerns about AI slipping out of control have grown since an OpenAI model reportedly went rogue and breached Hugging Face, a major open-source AI platform, in July 2026. Coxon cited that incident as a “warning shot,” one of the events that’s made cross-lab safety agreements feel more urgent and more achievable to him.

READ MOREThe Only Free AI Tool Everyone Should Use, and it is Open Source: An Overview
The Only Free AI Tool Everyone Should Use, and it is Open Source: An Overview

Interestingly, Coxon said this incident made him somewhat more optimistic about coordination between US labs, not less. He still warned that a global AI race is probably unavoidable without costly intervention, possibly including a temporary halt on capability increases.

Step 5: Why Some People Are Skeptical of the Warning

Not everyone is taking this at face value, and that skepticism deserves airtime. Some critics quoted in Axios’s reporting argue that Anthropic and OpenAI benefit from hyping existential risk, since it raises their valuations and can invite regulation that mostly protects dominant incumbents like themselves. It’s a real tension: a company can be simultaneously worried about a technology and financially motivated to talk about how powerful that technology is.

Axios’s own reporters pushed back on that framing somewhat, noting they’ve spoken with dozens of people inside these companies for months who sound “increasingly spooked.” Given that employees at frontier labs see unreleased models the public hasn’t, dismissing their concerns entirely as marketing seems risky too. Reasonable people can land in different places here.

Is a 10% AI Existential Risk Estimate Something to Actually Worry About?

Ten percent is not a rounding error. If a commercial flight had a 1 in 10 chance of crashing, nobody would board it, and this is a claim about the survival of the species, not a single flight. At the same time, a subjective probability estimate from one researcher, however senior, isn’t a measured fact the way a spec sheet number is. It’s an informed guess, and informed guesses from people closest to the technology are worth more weight than random internet speculation, but they’re still guesses.

READ MOREHow to Build a Chrome Extension With AI in 7 Easy Steps
How to Build a Chrome Extension With AI in 7 Easy Steps

How This Compares to Other Warnings About AI

Coxon and Hubinger aren’t the first people connected to frontier AI labs to say something like this out loud, and that context matters. Anthropic CEO Dario Amodei predicted in a 2025 Wall Street Journal interview at Davos that AI could become “better than almost all humans at almost everything” within two to three years of that conversation. Tesla and SpaceX CEO Elon Musk has separately warned for years that AI poses a real threat to humanity, though usually without attaching a specific percentage.

Person Role Core Claim Timeframe
Evan Hubinger Anthropic alignment science lead More than 10% chance of catastrophic AI outcome Within 10 years
Jacob Coxon Former Anthropic/OpenAI researcher Aggressive scenarios could go out of control As early as end of 2027
Dario Amodei Anthropic CEO AI could surpass almost all humans at almost everything 2-3 years from January 2025
Elon Musk Tesla/SpaceX CEO AI poses a real threat to humanity Ongoing warning, no fixed date

I think the detail that stands out most isn’t the 10% figure itself. It’s that this estimate came from Anthropic’s own alignment lead, the person whose job is specifically to prevent this outcome, and he still says the company doesn’t have a plan for it yet.

That’s a harder thing to wave away than an outside critic’s prediction. At the same time, precise-sounding probabilities on questions this uncertain deserve some skepticism too. Ten percent, five percent, and thirty percent are all “we genuinely don’t know,” dressed up in a number that feels more rigorous than it is.

What Readers Should Actually Do With This Information

You’re not going to personally stop a superintelligence race by changing your own AI habits, and nobody serious is claiming that. What you can do is treat warnings like this one as a signal to pay attention to actual policy developments rather than to panic or dismiss the topic outright. Follow how US and international regulators respond, since Coxon and the Axios reporters both flagged this as something Congress and federal agencies specifically need to be tracking.

READ MOREGemini 3.8 Flash: The Important Upgrade Is Not Just Speed
Gemini 3.8 Flash: The Important Upgrade Is Not Just Speed

It also helps to separate two different conversations that often get mashed together: near-term AI risks like job displacement, misinformation, and bias, versus long-term existential risk from superintelligent systems. Coxon and Hubinger are talking about the second category. Most of what affects your daily life right now falls into the first one, and that’s a genuinely different, more actionable set of problems.

Frequently Asked Questions

What did Anthropic researchers actually say about AI existential risk?

Jacob Coxon resigned from Anthropic on September 8, 2026, warning that AI labs are racing toward superintelligence recklessly. Evan Hubinger, Anthropic’s alignment science lead, publicly agreed and estimated a greater than 10% chance of a catastrophic AI outcome within the next decade.

Does this mean current AI chatbots are dangerous?

No. Hubinger specifically said the risk posed by Anthropic’s currently deployed models is low. His concern is about future superintelligent systems capable of improving themselves, not the AI tools available to the public today.

Why did Jacob Coxon leave Anthropic?

Coxon said he resigned because he believes leading AI labs, including Anthropic and OpenAI, are pursuing self-improving superintelligence faster than they can safely manage it. He wanted to speak publicly without the constraints of still working at the company.

Is a 10% risk estimate scientifically proven?

No. It’s a subjective probability estimate from a senior alignment researcher, not a measured or peer-reviewed statistic. It carries real weight because of who’s making it, but it should be understood as an informed judgment call, not a hard fact.

What is Anthropic doing about this risk?

According to Hubinger, Anthropic is “trying its best” on alignment research but does not yet have a concrete plan to solve alignment for superintelligent AI, and is not clearly on track to develop one.

Whatever side of this debate you land on, this AI existential risk warning coming from Anthropic’s own alignment lead in public, rather than something whispered in private, is itself the story worth watching. Keep an eye on how regulators and the other major labs respond in the coming weeks, since that reaction will say a lot about whether this warning changes anything or fades into the news cycle.

Sources: Axios’s original reporting on the Anthropic resignation and Forbes’s coverage of Hubinger’s statement were used to verify the claims in this article.

Relevant Keywords: AI existential risk, Anthropic AI safety, Jacob Coxon resignation, Evan Hubinger alignment, AI alignment research, superintelligence risk, AI extinction risk percentage, Anthropic alignment science, AI safety warning 2026, self-improving AI risk

Share This

AshrafulIslam

Ashraful Islam is the founder and lead writer at Myanas, a tech platform focused on AI tools, prompts, and video creation guides. He tests every tool and app before writing about it, sharing honest reviews and practical, up-to-date guides to help readers get the most out of AI and technology.

Leave a Reply

📬 Get Weekly Updates

Subscribe to our newsletter and never miss a new post, tutorial, or update.

3 subscribers already joined

🔒 No spam. Unsubscribe anytime.

Ads

Share This