OpenAI Withdrew 3 of Its 722 Math Papers Within 36 Hours. Tao Says the Whole Game Needs Rewriting.
Thirty-six hours after OpenAI dropped 722 machine-written math manuscripts on GitHub, the repository's history file got its first entry. Three papers withdrawn. A sign error in "Algebraicity of Weil classes on split abelian eightfolds" broke a stabilization-trace cancellation argument, and because two other manuscripts built on that construction, all three went down together, including a claimed proof of the rational Hodge conjecture for products of K3 surfaces. Fourteen more manuscripts got proof repairs, corrected statements and tightened hypotheses, and 13 others had their citations bumped to the revised editions. OpenAI's Daniel Litt announced it on X as "6 new Lean formalizations, 19 modifications, and 3 withdrawals," and the formalized share of top-line results is now 300 of 719, about 42%. Hacker News ran two threads on it, 337 and 234 points.
Nobody should be shocked that a sign error exists in 722 papers written in weeks. What is new is the speed and the visibility. The error was found by a reader, the withdrawal is public, the retracted manuscript is archived with a notice, and the dependent papers were pulled rather than left dangling. That is better practice than most human journals manage. It is also the first hard number on the verification backlog: 3 of 719 wrong inside two days, with 58% of results still unformalized and an untold number of readers who have not started.
Terence Tao's response landed the same morning, and it is the piece people will quote for years. In a four-part thread on Mathstodon that drew 585 points on Hacker News, Tao argues that "Math 1.0" put a premium on being first to crack an open problem even when the solution was not understood, and that this goal "has now been optimized to the point of unsustainability." "Math 2.0," in Tao's framing, has to decenter raw problem solving and value exposition, community building and opening new directions. AI can help with all of that, Tao says, but it takes "more imagination and ambition than the Math 1.0 mindset of simply pointing one's favorite AI agent at some set of open problems and asking for a solution." Education, publication and career advancement criteria all need rewriting.
The institutional reaction split hard. TechCrunch checked OpenAI's release against the Advisory Group on Mathematics and AI guidelines OpenAI said it was following: the group's first request was to stop testing open problems on proprietary models, which OpenAI explicitly did, only 10 of 719 manuscripts include the model's reasoning trace, and the group says it is "ultimately up to the mathematical community" to judge whether its recommendations were met. A newer group called the Association for Human Mathematics went further, calling the release "a demonstration of power, not of scholarship" and urging mathematicians to stop working with OpenAI. And on the crypto side, Ethereum's Justin Drake used the moment to push "bunker mode," on the theory that a system producing this much number theory this fast might eventually reach the assumptions under public key cryptography.
The honest read is that the output side of AI mathematics is now solved enough to be a problem. The binding constraint has moved to reading, checking and understanding, and that constraint is human time. Three retractions in 36 hours is not a scandal. It is what the backlog looks like when it starts to clear.
Links: github.com/openai/math/blob/main/history.md, mathstodon.xyz/@tao/117395269325940185, ahmath.org/statements, techcrunch.com/2026/10/08/openais-math-solutions-arent-meeting-the-fields-standards-yet/
← Back to all articles
Nobody should be shocked that a sign error exists in 722 papers written in weeks. What is new is the speed and the visibility. The error was found by a reader, the withdrawal is public, the retracted manuscript is archived with a notice, and the dependent papers were pulled rather than left dangling. That is better practice than most human journals manage. It is also the first hard number on the verification backlog: 3 of 719 wrong inside two days, with 58% of results still unformalized and an untold number of readers who have not started.
Terence Tao's response landed the same morning, and it is the piece people will quote for years. In a four-part thread on Mathstodon that drew 585 points on Hacker News, Tao argues that "Math 1.0" put a premium on being first to crack an open problem even when the solution was not understood, and that this goal "has now been optimized to the point of unsustainability." "Math 2.0," in Tao's framing, has to decenter raw problem solving and value exposition, community building and opening new directions. AI can help with all of that, Tao says, but it takes "more imagination and ambition than the Math 1.0 mindset of simply pointing one's favorite AI agent at some set of open problems and asking for a solution." Education, publication and career advancement criteria all need rewriting.
The institutional reaction split hard. TechCrunch checked OpenAI's release against the Advisory Group on Mathematics and AI guidelines OpenAI said it was following: the group's first request was to stop testing open problems on proprietary models, which OpenAI explicitly did, only 10 of 719 manuscripts include the model's reasoning trace, and the group says it is "ultimately up to the mathematical community" to judge whether its recommendations were met. A newer group called the Association for Human Mathematics went further, calling the release "a demonstration of power, not of scholarship" and urging mathematicians to stop working with OpenAI. And on the crypto side, Ethereum's Justin Drake used the moment to push "bunker mode," on the theory that a system producing this much number theory this fast might eventually reach the assumptions under public key cryptography.
The honest read is that the output side of AI mathematics is now solved enough to be a problem. The binding constraint has moved to reading, checking and understanding, and that constraint is human time. Three retractions in 36 hours is not a scandal. It is what the backlog looks like when it starts to clear.
Links: github.com/openai/math/blob/main/history.md, mathstodon.xyz/@tao/117395269325940185, ahmath.org/statements, techcrunch.com/2026/10/08/openais-math-solutions-arent-meeting-the-fields-standards-yet/
Comments