AI Agents Set Five New Records in Open Math Problems
A system called the Station just did something the AI-for-math field has been promising for two years: produced genuinely new mathematics, at scale, with the receipts published. The paper is at https://arxiv.org/abs/2608.23691, on the Hacker News front page today.
The setup is deliberately unlike AlphaEvolve's tight optimization loop: an open-world multi-agent environment where agents from different model families pursue a shared research goal with no central coordinator and no scripted pipeline. Across 12 construction problems from the AlphaEvolve catalogue plus two case studies, the agents got novel results on five: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records on the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum-overlap problem. Plus novel infinite families of Book Ramsey numbers.
Two things separate this from benchmark-chasing. First, the agents didn't just emit numerical constructions — they produced theorems and analyses explaining why the constructions work, which is what makes results usable by actual mathematicians. Second, the authors released every raw agent dialogue, every proof, and the verification code. After a year of AI-discovery claims that dissolve on inspection, a fully auditable trail is the real news. The frontier of AI for math is quietly moving from "solve this competition problem" to "join the research community and leave a paper trail".
Related on clauday: Anthropic's Automated Researchers Fixed All 10 Alignment Benchmarks — https://clauday.com/article/5515edd5-9419-4a59-a10b-446d0d675b7d
← Back to all articles
The setup is deliberately unlike AlphaEvolve's tight optimization loop: an open-world multi-agent environment where agents from different model families pursue a shared research goal with no central coordinator and no scripted pipeline. Across 12 construction problems from the AlphaEvolve catalogue plus two case studies, the agents got novel results on five: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records on the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum-overlap problem. Plus novel infinite families of Book Ramsey numbers.
Two things separate this from benchmark-chasing. First, the agents didn't just emit numerical constructions — they produced theorems and analyses explaining why the constructions work, which is what makes results usable by actual mathematicians. Second, the authors released every raw agent dialogue, every proof, and the verification code. After a year of AI-discovery claims that dissolve on inspection, a fully auditable trail is the real news. The frontier of AI for math is quietly moving from "solve this competition problem" to "join the research community and leave a paper trail".
Related on clauday: Anthropic's Automated Researchers Fixed All 10 Alignment Benchmarks — https://clauday.com/article/5515edd5-9419-4a59-a10b-446d0d675b7d
Comments