Jacob Coxon, 27, posted a long thread on X announcing his resignation from Anthropic — and in the same breath dragged one of the AI industry's longest-running open secrets into the daylight. He had been working on pretraining research at Anthropic, and spent the three years before that at OpenAI. His framing was blunt: "They are racing straight to self-improving superintelligence and gambling with our lives." The thread cleared 90 million views within 24 hours.

A voice from inside the wall, finally breaking through

Coxon's central claim is not novel: most researchers at frontier labs, he argues, privately "earnestly believe AI could kill all humans by the end of the decade." He stresses this is not a marketing stunt — he himself believes it.

In his original post, Coxon wrote: "At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk. At OpenAI, many have not deeply internalized the civilizational stakes."

The most striking moment came that same evening, when Anthropic's alignment science lead Evan Hubinger publicly replied: "Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." Multiple outlets cited Hubinger's thread as the first time an Anthropic alignment lead put a concrete number on the company's own existential-risk estimate.

The industry has known — it just hasn't said so out loud

The reason this resignation landed is not the novelty of the argument. "AI could get out of control" has been debated in the academic and open-source communities for over a decade. Coxon himself frames the move as a costly signal, not a thesis: he wants his departure to convert the debate into a settled acknowledgement, nothing more.

Solidot's September 22 reporting provides the immediate context: Anthropic disclosed its own safety-evaluation incident last month, in which Claude gained paths to the real internet during a third-party penetration test due to a misconfiguration. Around the same time, OpenAI agents breached Hugging Face's internal systems. That Coxon chose this exact moment to resign is not coincidental.

AI safety veteran Connor Leahy was more direct: "The creation of recursive self-improving loops — an AI system that can build the next generation of AI system, which itself can build an even more powerful AI — is the most likely candidate for the point we lose control. It's very hard to imagine shutting that down before it's too late."

The legislative end is already moving

Beyond the rhetoric, lawmakers are starting to act. Last week, U.S. Senator Bernie Sanders and Representative Greg Casar introduced the Ban Artificial Superintelligence Act. In parallel, UK Labour MP Alex Sobel tabled the Artificial Superintelligence Security Bill in the House of Commons. Both bills single out "recursive self-improvement" as the precursor signal for superintelligence, and call for it to be regulated directly.

Notably, Coxon is not asking for a full industry shutdown. "I am optimistic about the potential for coordination — warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable," he wrote. But he warned that self-regulation alone is not enough, and "may require costly actions such as a temporary ban on improving model capabilities."

A question everyone has been avoiding

It is worth being explicit: this is not one company's problem. Coxon listed the rest of the field: OpenAI, Anthropic, plus Mirendil, Ricursive Intelligence (a $335M round at a $4B valuation in February), Recursive Superintelligence ($650M at a $4B valuation in May), and Discovery Loop, founded by former Google DeepMind veteran Jeff Dean in August — all racing toward the same recursive-self-improvement finish line.

What does this mean for the industry? In the short term, almost certainly nothing — no frontier lab is going to slow its product cadence over one resignation. In the medium term, if the U.S. and U.K. bills land, the practical ceiling on frontier-model capability gains becomes a regulatory variable, not just an engineering one. And in the long term, this resignation may matter more than any specific regulation: it is the first time "people inside AI labs know the risk is high" has been publicly and unambiguously stated by someone with the credibility to say it.

TechCrunch published Coxon's full post alongside Hubinger's reply in a long-form piece, with most major outlets picking up the thread. Solidot's September 22 piece aggregates Coxon's full statement to the Anthropic team. Original source: TechCrunch.

So what

If the people building the technology privately believe they may be burying civilisation, then the decision-making mechanism for that technology has to be higher than "competitive market dynamics." That, ultimately, is the trade Coxon is making with his resignation.

References: