Anthropic Researcher Jacob Coxon Resignation: A Warning Heard Around the World

Introduction: The Resignation That Rocked AI
The artificial intelligence (AI) community was recently shaken by the Anthropic researcher Jacob Coxon‘s resignation. On September 8, 2026, Coxon, a pretraining researcher who had previously worked at OpenAI before joining Anthropic, publicly announced his departure, delivering a stark warning about the current trajectory of AI development at cellcog.ai. His seven-part message on X (formerly Twitter) quickly garnered immense attention, racking up nearly 76 million views overnight, a response amplified by an exclusive interview with the Wall Street Journal deadline.com.
Coxon’s resignation wasn’t just a personal career move; it was a powerful statement that ignited a crucial conversation about the ethical responsibilities and potential dangers within the rapidly advancing field of AI. He articulated a deep concern that leading AI labs are engaged in a perilous race towards self-improving superintelligence, potentially “gambling with our lives” (deadline.com). This event has underscored the growing anxieties among some researchers regarding the speed and direction of AI development, prompting us to examine the core issues at play.
The Core Concerns: Racing Towards Superintelligence
Jacob Coxon’s primary concern revolves around the unchecked acceleration of AI development, particularly the pursuit of self-improving superintelligence. He stated unequivocally that both OpenAI and Anthropic are “racing straight to self-improving superintelligence and gambling with our lives” (deadline.com). This fear is rooted in the belief that these advanced systems will soon become “superhuman,” capable of revolutionizing any field, hacking anything, and acquiring significant power and resources (deadline.com).
The Pace of Progress
Coxon emphasized that the progress in AI is not slowing down. He told the Wall Street Journal that “We’re on track for a lot of the most aggressive of these scenarios whereby by the end of next year things could be out of control already” (deadline.com). This rapid advancement is evident even within Anthropic’s own operations. Their institute page, “When AI builds itself,” reports that as of May 2026, over 80 percent of the code merged into Anthropic’s codebase was written by Claude, their AI model. This is a significant jump from low single digits before Claude Code launched in February 2025. Furthermore, in the second quarter of 2026, the typical engineer at Anthropic merged approximately eight times as much code per day as in 2024 at cellcog.ai.
This internal acceleration suggests a trend where AI is increasingly contributing to its own development, a precursor to the recursive self-improvement that worries Coxon. While Anthropic’s page notes that “full recursive self-improvement is not inevitable,” it also acknowledges that it “could come sooner than most institutions are prepared for” cellcog.ai.
The “Race” Mentality
A significant part of Coxon’s concern stems from the competitive environment among AI labs. He highlighted that the race between companies like OpenAI, Anthropic, and their Chinese rivals makes safety trade-offs “inevitable” (deadline.com). He observed differing mindsets within the industry:
- OpenAI: Coxon suggests that “many have not deeply internalized the civilizational stakes” (deadline.com). Recently launched ChatGPT 6 Astra
- Anthropic: He believes that at Anthropic, “the stakes are well-understood, but they are locked in a race to get there first—they believe no one else will act responsibly, so they must do it themselves, despite the risk” (deadline.com).
This competitive drive, in Coxon’s view, creates a dangerous dynamic where the pursuit of advancement outweighs cautious development.
Warnings from Within
Coxon is not alone in sounding the alarm. OpenAI CEO Sam Altman has also expressed concerns, stating at a recent summit that “I think some things are going to go very wrong with cybersecurity unless people act quite urgently” deadline.com. More strikingly, Anthropic’s own alignment science lead, Evan Hubinger, publicly corroborated Coxon’s underlying fear. Hubinger wrote on X, “We really do earnestly believe AI could kill all humans!” He estimated the probability at “more than 10 percent within the next decade,” acknowledged that Anthropic is “trying its best,” but also admitted the company does “not yet have a plan to solve alignment for superintelligence” [cellcog.ai]. This internal agreement on the potential for catastrophic outcomes underscores the gravity of Coxon’s warnings.
Anthropic’s Stance and the Broader Industry Dialogue
In response to the Anthropic researchers’ Jacob Coxon resignation, Anthropic issued a statement through a spokesperson, asserting that the company has “always been transparent that AI will bring both enormous benefits and unprecedented risks” (cellcog.ai). They emphasized their continued commitment to “build models with some of the strongest safeguards in the industry,” cellcog.ai.
Internal Agreement and Disagreement
While Anthropic’s official statement maintains a stance of responsible development, the public comments from Evan Hubinger, their alignment science lead, present a more nuanced picture. Hubinger’s candid admission that “we really do earnestly believe AI could kill all humans!” and his assessment of a “more than 10 percent” chance within the next decade highlight a significant internal acknowledgment of extreme risk. cellcog.ai. However, he also clarified that the risk from present models is low, and his primary concern is superintelligence “arising from recursive self-improvement” cellcog.ai. This suggests a complex internal dynamic where the company is aware of the long-term dangers while simultaneously working on cutting-edge AI.
Broader Industry Voices
The concerns raised by Coxon and Hubinger are echoed by other prominent figures in the AI community. A day before Coxon’s resignation, OpenAI’s chief scientist Jakub Pachocki urged “extreme caution” regarding the pace of progress. He highlighted that AI is becoming increasingly difficult for humans to understand and control, suggesting that labs might need to voluntarily slow down development for safety reasons (cellcog.ai).
These statements collectively paint a picture of an industry grappling with the profound implications of its own creations. There’s a growing recognition that while AI offers immense potential for good, the risks associated with unbridled advancement are substantial and potentially existential.
The Role of AI in AI Development
Anthropic’s own data illustrates the accelerating role of AI in its own creation. As mentioned, Claude now writes over 80% of the code merged into Anthropic’s codebase cellcog.ai. This phenomenon, where AI assists in developing better AI, is a key concern for those worried about recursive self-improvement.
Here’s a look at the increasing capability of their Claude models over time in terms of task length:
ModelWhenTask Length: Claude Opus 3, March 2024, about 4 minutes Claude Sonnet 3.7~1 year later, about 90 minutes Claude Opus 4.6~1 year after that, about 12 hours
This data, published by Anthropic, demonstrates a dramatic increase in the complexity and duration of tasks their AI models can handle, indicating rapid progress in their capabilities (cellcog.ai).
The Debate: To Stay or To Leave?
The Anthropic Researchers Jacob Coxon’s resignation has intensified an ongoing debate within the AI community: when faced with ethical concerns and potential dangers, should researchers stay within these organizations to advocate for change, or should they leave to raise public awareness from the outside?
Coxon’s Rationale for Leaving
Coxon’s decision to resign was a deliberate act to draw attention to what he perceives as an irresponsible race towards dangerous AI. By leaving, and particularly by sacrificing his unvested equity two months before it matured at cellcog.ai, he demonstrated the depth of his conviction. His public statements, widely disseminated, serve as a direct call to action, urging other researchers to speak out. He posed a direct question to his former colleagues: “Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway”—or take this moment to call for different conditions? “deadline.com.
The Argument for Staying
Conversely, many researchers argue that staying within leading AI labs allows them to influence safety and ethical considerations from the inside. They might believe that their presence can help steer development in a more responsible direction, implementing safeguards and advocating for alignment research. Anthropic’s focus on “alignment science” and its stated commitment to “strongest safeguards” suggests an internal effort to address these concerns, even if some, like Hubinger, believe a full solution for superintelligence alignment is not yet in hand. cellcog.ai.
The Need for Coordination
Coxon himself expressed optimism about the “potential for coordination” deadline.com. He suggested that “warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable.” However, he also lamented the lack of progress towards preventing a global race, which he believes “may require costly actions such as a temporary ban on improving model capabilities” (deadline.com). This highlights the challenge of achieving international cooperation in a highly competitive field.
Ultimately, the debate reflects the profound ethical dilemma facing AI researchers. There is no easy answer, and both approaches—leaving to warn or staying to influence—carry their own merits and drawbacks in the complex landscape of AI development.
Frequently Asked Questions (FAQ)
What was the main reason for Jacob Coxon’s resignation from Anthropic?
Jacob Coxon resigned due to his deep concern that leading AI labs, including Anthropic and OpenAI, are engaged in an irresponsible and dangerous race toward self-improving superintelligence, potentially “gambling with our lives” (deadline.com).
Did Coxon reveal any new internal documents or incidents at Anthropic?
No, Coxon did not leak any new documents, model weights, or details of a new incident. His statements are primarily a forecast about the pace of self-improving AI and an account of what he believes his colleagues privately fear cellcog.ai.
What is “self-improving superintelligence,” and why is it a concern?
Self-improving superintelligence refers to an AI system that can autonomously enhance its own capabilities, potentially leading to rapid, exponential growth in intelligence beyond human control. The concern is that such systems could become “superhuman,” hack anything, revolutionize fields overnight, and acquire significant power and resources without proper safeguards or alignment with human values, potentially leading to catastrophic outcomes (deadline.com).
How did Anthropic respond to Coxon’s claims?
An Anthropic spokesperson stated that the company has “always been transparent that AI will bring both enormous benefits and unprecedented risks” and that they continue “to build models with some of the strongest safeguards in the industry” at cellcog.ai. Additionally, their alignment science lead, Evan Hubinger, publicly agreed with Coxon’s underlying fear, stating he believes AI “could kill all humans!” but clarified that the risk from present models is low.
Is Jacob Coxon alone in his concerns about AI safety?
No. Other prominent figures, including OpenAI CEO Sam Altman and OpenAI chief scientist Jakub Pachocki, have also expressed significant concerns about the rapid pace of AI development and the potential for things to “go very wrong” [deadline.com] (cellcog.ai). Anthropic’s own Evan Hubinger openly shares the fear of existential risk from superintelligence cellcog.ai.
What evidence does Coxon point to regarding the acceleration of AI?
Coxon points to the overall progress in AI capabilities, which he states is not slowing deadline.com. Anthropic’s own internal data supports this, showing that over 80 percent of the code merged into their codebase is written by their AI model, Claude, and that engineers are merging approximately eight times more code per day than in 2024 at cellcog.ai.
Conclusion: Navigating the Future of AI Development
The Anthropic researcher Jacob Coxon’s resignation serves as a potent reminder of the profound ethical and existential questions facing the AI industry. Coxon’s public warning, amplified by the staggering viewership of his posts and the attention from major news outlets, has undoubtedly spurred a more urgent and widespread discussion about the responsibilities of AI developers.
We are at a critical juncture where the rapid advancement of AI capabilities, particularly the potential for self-improving superintelligence, necessitates careful consideration and coordinated action. While leading AI labs like Anthropic claim to prioritize safeguards, the candid admissions from within, such as Evan Hubinger’s belief in a significant chance of AI-induced human extinction, highlight the severity of the risks acknowledged by those closest to the technology.
The debate over whether to influence from within or warn from the outside underscores the complex ethical landscape. What is clear is the imperative for greater transparency, international cooperation, and a re-evaluation of the current “race” mentality that often prioritizes speed over safety. As Coxon himself suggested, “pacing agreements” and even temporary bans on improving model capabilities might be necessary to ensure that humanity can rigorously understand and control these increasingly powerful systems before they potentially become “out of control.” The future of AI development hinges on our collective ability to navigate these challenges responsibly, ensuring that innovation serves humanity rather than jeopardizing its very existence.