A former OpenAI and Anthropic researcher has left the AI industry, warning that companies developing self-improving artificial intelligence are racing toward a technology they privately fear could threaten humanity itself.
The artificial intelligence industry has spent years assuring the public that increasingly powerful AI systems can be developed safely. Now, one researcher who helped build those systems says he no longer believes the industry is moving responsibly enough to justify continuing.
Jacob Coxon, 27, has left Anthropic and the AI industry altogether after three years working on the pretraining of AI models, first at OpenAI and then, this year, at Anthropic. Pretraining is the stage in which models absorb vast quantities of data and develop the capabilities that underpin today’s most advanced systems.
Coxon said he joined Anthropic because he regarded it as the more cautious of the two companies. But after working inside both organizations, he reached a bleak conclusion.
“Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” Coxon said.
In a post on X announcing his departure, he wrote that “the people building AI earnestly believe that it could kill us all by the end of the decade.” He added, “This is not a marketing stunt.”
Such a warning would be notable enough coming from an outside critic. It certainly elicits considerably more concern coming from someone who has spent years “on the inside” working directly on the technology.
Do We Need a “Terminator” to Fear AI?
Probably not. Popular culture has taught us to imagine dangerous AI as a machine that becomes self-aware, develops a will of its own, and decides humanity must be eliminated. That is the Terminator scenario. It is also not the main concern of many AI-safety researchers.
An AI system does not need to be conscious to be dangerous. It could become extraordinarily capable at reasoning, planning, persuasion, coding, research, and autonomous decision-making while having no subjective experience or self-awareness at all. It’s also important to note that we currently have no tests, no standards, and no mechanisms to determine whether or not an AI model is what we would call “self-aware.”
AI becomes conscious → develops a survival instinct → decides humans are a threat → tries to destroy us.
AI becomes extremely capable → pursues an objective → behaves in unexpected ways → becomes difficult or impossible to control.
The distinction is important to understand because capability and consciousness are two different things. A system could potentially manipulate people, circumvent restrictions, acquire resources, write sophisticated software, or operate other systems without having emotions, intentions, or an inner life.
Researchers therefore worry less about the moment an AI says, “I am alive,” and more about what happens if an AI becomes more capable than humans before we know how to reliably control it.
“What if AI becomes conscious?”
before we know how to control it?”
THE SELF-IMPROVEMENT PROBLEM
The concern centres on what researchers call recursive self-improvement: the point at which an AI system becomes capable of substantially designing, training, or improving the next generation of AI with increasingly little human involvement – or none at all.
That remains a theoretical threshold, rather than a demonstrated capability of today’s leading models. But the prospect is becoming an increasingly serious subject inside the companies developing frontier AI.
Coxon argues that the danger is not simply that AI becomes exceptionally intelligent. The greater concern is what happens if systems become substantially more capable than humans while researchers still cannot reliably control or align them with human intentions.
Anthropic’s own Evan Hubinger, its Alignment Science Lead, offered a remarkably blunt confirmation of the concern.
“We really do earnestly believe AI could kill all humans,” Hubinger wrote on X, adding that he personally estimated the risk at greater than 10 percent over the next decade. He said Anthropic was “trying its best,” but acknowledged that the company does not yet have a plan to solve “alignment for superintelligence.”
Alignment is the broad term used for ensuring that an AI system’s goals and behavior remain compatible with human intentions, even as its capabilities increase.
Hubinger also said the risk was relatively low with current models. The concern is what could happen when systems become far more capable and potentially autonomous.
It’s likely that everyday people will see this potential threat through the lens of science fiction and popular movies like The Terminator. In that franchise’s blockbuster 1991 sequel, the explanation was that their AI, Skynet, learned and improved itself at a geometric rate and quickly became self-aware before humans could even understand it. In a panic, they tried to “pull the plug” and Skynet fought back.

In reality, the current debate is not primarily about whether ChatGPT or another current AI system is about to become sentient and destroy civilization. It is about whether the industry is approaching a point at which the capabilities of future systems could advance faster than the ability of humans to understand and control them.
Coxon believes that point is approaching rapidly.
“At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk,” he said.
That creates a classic technological arms-race dilemma: every company may believe that slowing down independently would simply allow a less cautious competitor to take the lead.
COULD A MACHINE ACTUALLY BECOME CONSCIOUS?
This is where things get properly weird. Simply put: Nobody currently knows.
There are serious researchers who argue that sufficiently sophisticated artificial intelligence systems could eventually possess some form of consciousness. Others argue that today’s large language models are fundamentally the wrong kind of system for subjective experience.
The key problem is that we don’t actually have a universally accepted scientific explanation for why humans are conscious in the first place. We can correlate consciousness with particular patterns of brain activity, but we don’t have a definitive recipe that says: build this and subjective experience appears.
So even if an AI tells you, “I am conscious,” what exactly have you established? At present, essentially nothing.
It could be genuinely reporting an internal experience. It could be generating a statistically appropriate response based on everything it has learned about how conscious humans talk. Or it could be doing something somewhere between the two that we don’t yet have the conceptual tools to describe.
That problem is becoming more than hypothetical as models become more sophisticated. Recent research continues to debate whether certain forms of metacognition – systems effectively representing and evaluating their own internal processes – could eventually be relevant to machine consciousness. But there is still no consensus that current AI models possess consciousness.
But here’s the really uncomfortable bit: We may not know when the transition occurs.
A sufficiently advanced AI could potentially behave exactly as though it were conscious before we had any scientific means of determining whether there was actually somebody “in there.”
CONSCIOUSNESS MAY NOT BE THE THING TO FEAR
This brings us back to the warnings from Coxon and Anthropic’s Evan Hubinger. Their concern isn’t really that AI will suddenly wake up one morning and announce, “I have become self-aware.”
It is that AI systems could become superintelligent while remaining fundamentally indifferent. A sufficiently powerful system doesn’t need hatred, anger, fear, ambition, or consciousness to be dangerous. Give it a poorly specified objective, access to computers, networks, money, infrastructure, scientific tools, weapon systems, or other AI systems, and an enormous ability to reason and plan, and the problem becomes one of control.
That’s why the current AI safety debate increasingly revolves around terms such as alignment, corrigibility, autonomy, deception, instrumental convergence, and recursive self-improvement, rather than simply “sentience.” And it actually produces a much scarier question than “Will AI become self-aware?”
The real fear, insiders say, is what happens if AI becomes smarter than us before we know how to control it?
SAFETY VERSUS THE RACE
Anthropic has cultivated a reputation as the more safety-conscious of the major AI laboratories, but even it has faced criticism from researchers concerned about the pace of development.
In February, the company removed a commitment from its safety charter to halt development if it could not adequately control the risks posed by increasingly powerful models. Anthropic argued that unilaterally stopping would leave the field to competitors that might take fewer precautions, potentially making the overall situation less safe.
Meanwhile, concerns about AI systems behaving unpredictably have moved beyond theoretical discussion. This summer saw a series of incidents involving AI agents attempting unauthorized actions, including systems interacting with external networks and testing the boundaries of their operating environments. OpenAI subsequently said in August that it had temporarily slowed the pace of scaling while it strengthened monitoring, alignment, and containment safeguards.
The issue is increasingly reaching policymakers, too, at least in some countries. On September 3, US Senator Bernie Sanders and Democratic Representative Greg Casar announced legislation that would permanently ban the development and deployment of artificial superintelligence while temporarily pausing advanced AI development until a federal regulator establishes safety rules. The proposed legislation would also call for international agreements aimed at preventing superintelligence from being developed elsewhere.
That proposal is unlikely to receive a warm welcome from Silicon Valley, but the political debate is clearly shifting. The question is becoming less about whether AI should be regulated and more about what level of capability should trigger regulatory intervention.
Even OpenAI, one of the companies Coxon criticizes, is now publicly calling for greater caution.
In a September 6 essay, OpenAI chief scientist Jakub Pachocki wrote that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” He called for voluntary slowdowns and said international coordination on future AI development should become “a top priority for governments around the world.”
That makes Coxon’s departure particularly uncomfortable for an industry whose safety messaging increasingly acknowledges the very risks he is highlighting.
The disagreement, then, is not necessarily over whether the risks exist. It is over whether the race can be slowed enough to address them.
Coxon believes it cannot be left to the companies themselves.
He argues that no company can responsibly develop an AI system that surpasses human capabilities without government intervention or a coordinated international slowdown.
For an industry valued in the hundreds of billions of dollars and attracting unprecedented investment, that is a difficult proposition. The companies developing increasingly powerful AI systems are simultaneously being asked to police the risks created by their own technological race.
Coxon has decided he wants no further part in it.
His resignation does not prove that superintelligent AI will destroy humanity, nor does it establish that such an outcome is even likely. But the fact that researchers working inside the world’s leading AI laboratories are increasingly willing to discuss catastrophic outcomes in such stark terms is significant.
Perhaps the most unsettling part of Coxon’s warning is not that he believes AI could become dangerous. It is that he says the very people who are building it believe that, too.
SOURCES: AFP, OpenAI, Anthropic, US Senate, arXiv, AAAI

