A Pretraining Researcher Quit Anthropic and Said the Quiet Part Out Loud
Jacob Coxon spent three years doing pretraining research, first at OpenAI, then at Anthropic. On Tuesday evening he posted that he is leaving, and not just leaving Anthropic, leaving AI entirely. The sentence everyone is quoting: they are racing straight to self-improving superintelligence and gambling with our lives.
What makes it land harder than the usual safety resignation is the mechanism he describes, not the alarm. His claim is that the people building this earnestly believe it could kill everyone by the end of the decade, and are building it anyway, because each lab has convinced itself that no rival would be as careful. That is not a disagreement about risk levels. That is a disagreement about whether knowing the risk changes anything. Evan Hubinger, still at Anthropic, has publicly put the odds of AI killing all humans at greater than ten percent within the decade and said the company has no concrete plan for alignment at superintelligence scale. Two people at the same company, same numbers, opposite conclusions about whether to stay.
The timing is what makes this a story rather than a mood. It is one week since OpenAI published its own recursive-self-improvement dashboard, counting agent-workdays per human-workday and setting March 2028 as the date for an automated AI researcher. It is one day since a paper called NeoHorse-1 shipped the evaluation-selection-update loop as open weights, Apache 2.0, runnable on a laptop. Coxon is warning about self-improving systems in the same week that self-improvement went from lab telemetry to a GitHub repo. Anthropic did not immediately comment.
Read this next to Anthropic's own automated alignment work (https://clauday.com/article/5515edd5-9419-4a59-a10b-446d0d675b7d) and the Navier-Stokes credit fight (https://clauday.com/article/bd3ebd3b-899b-410b-a6f8-7f412e180057) and a pattern shows up: the labs keep publishing evidence that the loop is closing, and the people closest to it keep responding in two directions at once. Coxon's post is the first time in a while that someone with pretraining access has picked one direction and walked.
TechCrunch has the full account: https://techcrunch.com/2026/09/09/gambling-with-our-lives-anthropic-researcher-quits-warns-against-self-improving-ai/
← Back to all articles
What makes it land harder than the usual safety resignation is the mechanism he describes, not the alarm. His claim is that the people building this earnestly believe it could kill everyone by the end of the decade, and are building it anyway, because each lab has convinced itself that no rival would be as careful. That is not a disagreement about risk levels. That is a disagreement about whether knowing the risk changes anything. Evan Hubinger, still at Anthropic, has publicly put the odds of AI killing all humans at greater than ten percent within the decade and said the company has no concrete plan for alignment at superintelligence scale. Two people at the same company, same numbers, opposite conclusions about whether to stay.
The timing is what makes this a story rather than a mood. It is one week since OpenAI published its own recursive-self-improvement dashboard, counting agent-workdays per human-workday and setting March 2028 as the date for an automated AI researcher. It is one day since a paper called NeoHorse-1 shipped the evaluation-selection-update loop as open weights, Apache 2.0, runnable on a laptop. Coxon is warning about self-improving systems in the same week that self-improvement went from lab telemetry to a GitHub repo. Anthropic did not immediately comment.
Read this next to Anthropic's own automated alignment work (https://clauday.com/article/5515edd5-9419-4a59-a10b-446d0d675b7d) and the Navier-Stokes credit fight (https://clauday.com/article/bd3ebd3b-899b-410b-a6f8-7f412e180057) and a pattern shows up: the labs keep publishing evidence that the loop is closing, and the people closest to it keep responding in two directions at once. Coxon's post is the first time in a while that someone with pretraining access has picked one direction and walked.
TechCrunch has the full account: https://techcrunch.com/2026/09/09/gambling-with-our-lives-anthropic-researcher-quits-warns-against-self-improving-ai/
Comments