OpenAI Dropped 722 Math Papers on GitHub. Nobody Has Read Them Yet.
On Tuesday evening OpenAI pushed a GitHub repository called openai/math and Hacker News gave it 1,205 points by morning. Inside are 722 manuscripts in 372 families, produced by an unreleased internal model, covering results that working mathematicians had chased for decades. A proof of the Unique Games Conjecture. L=BPL. Integer multiplication below n log n, breaking a bound from the 1960s. Matrix multiplication in n to the 9/4. A counterexample to Kaplansky's direct-finiteness conjecture. Counterexamples to Ryser's covering conjecture and Kurosh's division ring problem. Scott Aaronson, whose wife Dana Moshkovitz spent a career on the Unique Games Conjecture, called it "surely one of the biggest days in mathematical history" and then quoted Moshkovitz's first-night reaction: the paper "feels like something written by someone who's on psychedelics" and is "impossible to read without AI help."
The method is the interesting part for anyone building agents. OpenAI posed roughly 4,000 problems to one internal model. Each result averaged three hours of ChatGPT Pro thinking compute. The output was then aggregated into families and filtered for significance. That is the whole procedure, and it produced what looks like a decade of theory output in weeks. Ten reasoning summaries are published alongside, including the irrationality exponent of pi, the Mezard-Parisi formula for diluted spin glasses, and the three-dimensional relativistic Vlasov-Maxwell system. The repository is the response to the Advisory Group on Mathematics and AI at the Institute for Advanced Study, which asked labs on September 29 to publish with citations, revisions and versioning. OpenAI says it will fund workshops around understanding the results and is "working to responsibly release the model."
Now the caveats, which are large. OpenAI's own README says the collection "includes results at different stages of verification," that "not all have accompanying Lean formalizations," and that "some of the unformalized results could have issues." Moshkovitz found citations that are "often irrelevant and confusing" and a construction Moshkovitz describes as "some alien craziness." And on the same day, a paper from Bastounis, Circelli and Hansen (arXiv 2610.08144, HN 191) argues that a Lean certificate does not certify the natural-language proof it was translated from. They show that resolving ambiguity in mathematical prose faithfully sits arbitrarily high in the Solvability Complexity Index hierarchy, harder than the halting problem, and they claim that OpenAI's earlier Navier-Stokes Lean proof does not correspond to the natural-language argument it was supposed to verify. If that holds, the verification story around these 722 papers is thinner than the Lean badge suggests.
The honest framing is that the race to understand just started, and the understanding is going to be done with AI help because the papers are unreadable without it. Aaronson's prediction is a math world that is "heavenly if you have vision or creative ideas that AI could help check and implement." The near-term reality is 372 families of claims, a subset formalized, a reader base of maybe a few hundred people per family, and a translation problem that a theory paper says is uncomputable in the worst case. That is a verification backlog, not a finish line.
Links: github.com/openai/math, openai.com/index/sharing-ai-progress-in-mathematics/, scottaaronson.blog/?p=10169, arxiv.org/abs/2610.08144
← Back to all articles
The method is the interesting part for anyone building agents. OpenAI posed roughly 4,000 problems to one internal model. Each result averaged three hours of ChatGPT Pro thinking compute. The output was then aggregated into families and filtered for significance. That is the whole procedure, and it produced what looks like a decade of theory output in weeks. Ten reasoning summaries are published alongside, including the irrationality exponent of pi, the Mezard-Parisi formula for diluted spin glasses, and the three-dimensional relativistic Vlasov-Maxwell system. The repository is the response to the Advisory Group on Mathematics and AI at the Institute for Advanced Study, which asked labs on September 29 to publish with citations, revisions and versioning. OpenAI says it will fund workshops around understanding the results and is "working to responsibly release the model."
Now the caveats, which are large. OpenAI's own README says the collection "includes results at different stages of verification," that "not all have accompanying Lean formalizations," and that "some of the unformalized results could have issues." Moshkovitz found citations that are "often irrelevant and confusing" and a construction Moshkovitz describes as "some alien craziness." And on the same day, a paper from Bastounis, Circelli and Hansen (arXiv 2610.08144, HN 191) argues that a Lean certificate does not certify the natural-language proof it was translated from. They show that resolving ambiguity in mathematical prose faithfully sits arbitrarily high in the Solvability Complexity Index hierarchy, harder than the halting problem, and they claim that OpenAI's earlier Navier-Stokes Lean proof does not correspond to the natural-language argument it was supposed to verify. If that holds, the verification story around these 722 papers is thinner than the Lean badge suggests.
The honest framing is that the race to understand just started, and the understanding is going to be done with AI help because the papers are unreadable without it. Aaronson's prediction is a math world that is "heavenly if you have vision or creative ideas that AI could help check and implement." The near-term reality is 372 families of claims, a subset formalized, a reader base of maybe a few hundred people per family, and a translation problem that a theory paper says is uncomputable in the worst case. That is a verification backlog, not a finish line.
Links: github.com/openai/math, openai.com/index/sharing-ai-progress-in-mathematics/, scottaaronson.blog/?p=10169, arxiv.org/abs/2610.08144
Comments