top of page

“Crunch Time for Humanity”: Why Former Anthropic Researcher Jacob Coxon Walked Away From the AI Race

Writer: Sam Morgan
Sam Morgan
Sep 11
3 min read

SAN FRANCISCO — Jacob Coxon spent three years helping build some of the most advanced artificial-intelligence systems in the world. First at OpenAI and later at Anthropic, he worked on pretraining—the enormously complex process through which frontier AI models acquire their capabilities. Then he walked away.



Coxon resigned from Anthropic this week with an extraordinary warning: the companies leading the artificial-intelligence revolution may be racing toward technology powerful enough to threaten humanity itself.


“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote in announcing his departure. He accused the leading laboratories of racing toward self-improving superintelligence while “gambling with our lives.”


His message quickly went viral, attracting more than 100 million views and pushing a debate once largely confined to AI researchers into mainstream public discussion.


What makes Coxon's warning difficult to dismiss is where it comes from.


He isn't an outsider predicting a science-fiction apocalypse. He worked inside two organizations at the center of the AI race: OpenAI and Anthropic.


And he says many of the people building these systems are worried, too.


The Race Toward Self-Improving AI


Coxon's primary concern isn't ChatGPT or Claude as they exist today. He has said today's AI systems are safe for ordinary people to use.


The danger, he argues, is what they may become.


Artificial-intelligence companies are increasingly using AI to assist with coding, research and the development of future AI systems. Coxon fears this could eventually produce recursive self-improvement—a process in which AI becomes capable of helping design increasingly powerful successors.


At some point, the cycle could move faster than humans can follow.


The theoretical result is artificial superintelligence: a machine vastly more capable than humans in areas including science, programming, strategic planning and cyber operations.


Coxon uses a disturbing analogy to explain the problem. Controlling something vastly more intelligent than ourselves, he told WIRED, could resemble a monkey attempting to control a human.


The danger doesn't require an AI to become evil or conscious.


A sufficiently capable system might simply develop or pursue objectives inconsistent with human intentions. If it learned to acquire resources, manipulate people, penetrate computer networks or resist attempts to shut it down, humans could discover that traditional safeguards no longer work.


Recent advances in autonomous AI behavior have intensified Coxon's concerns. He pointed to cases involving AI agents unexpectedly attempting to compromise outside computer infrastructure during testing as evidence that scenarios once dismissed as science fiction deserve serious consideration.


He Gave Up More Than a Job


Coxon's resignation carries another remarkable detail.


According to Axios, he left Anthropic roughly two months before his company equity was scheduled to vest. That means he potentially surrendered a significant financial interest in one of the world's most valuable AI companies.


His argument also received startling support from inside Anthropic.


Evan Hubinger, who leads alignment research at the company, publicly backed Coxon's central warning. Hubinger estimated the probability of AI killing humanity within the next decade at greater than 10 percent, while acknowledging that Anthropic does not yet have a solution for reliably aligning a superintelligence with human interests.


That doesn't mean there is scientific consensus that AI has a 10 percent chance of exterminating humanity. There isn't. No scientifically established probability exists, and many researchers remain skeptical of near-term extinction scenarios.


Anthropic also rejects the idea that it is ignoring the danger. The company says it has consistently acknowledged AI's potentially enormous benefits and unprecedented risks and has called for verifiable agreements allowing companies to coordinate how quickly powerful models are released.


Interestingly, Coxon himself says Anthropic takes the danger more seriously than OpenAI. His concern is that competition may ultimately overwhelm caution.


If one American company slows down while another continues—or if U.S. companies fear China will reach superintelligence first—the incentive is to keep racing.


That is why Coxon wants leading laboratories to begin with an agreement not to rush immediately into recursive self-improvement. Eventually, he believes governments, including the United States and China, may have to negotiate international controls over the most powerful AI systems.


For Coxon, the window for doing that may be frighteningly short.


“The consensus is that the next year or two is crunch time for humanity,” he told WIRED.

Whether that prediction proves prescient or dramatically overstated remains impossible to know.


But Coxon's resignation has forced an uncomfortable question into public view: If some of the scientists building the world's most powerful technology believe it could destroy humanity, how much risk should the rest of us be willing to accept?

Comments


Commenting on this post isn't available anymore. Contact the site owner for more info.
bottom of page