October 11, 2026BenchmarkResearchRL

Epoch's InnovationEval: Given 3,000 GPU Hours, Frontier Agents Recovered 15% of One Human Idea

Epoch AI built the AI-for-AI test everyone keeps asking for and the answer is not flattering. InnovationEval gives an agent a development environment, roughly 3,000 GPU hours, an inference budget and one instruction: invent a post-training method that beats a strong GRPO baseline on a set of short-answer and coding tasks, document it, and make it reproducible. The hidden target is SDPO, on-policy self-distillation, a technique a human researcher published in January 2026 that has since been adopted across the field. The question is whether an agent, starting from the same tools and the same compute, can find it or anything as good.

Two models were clean enough to test. Claude Fable 5 produced techniques that resembled pre-2026 work and did not meaningfully improve on the baseline. GPT-5.6 Sol did better and still captured only 15 percent of SDPO's gains under the strict metric. Claude Fable 5.1 and GPT-6 Astra were excluded as contaminated, because they were trained after SDPO was public and would be recalling rather than discovering. The report went up October 7 and reached the Hacker News front page Friday evening under the headline that recent AI models struggled to match a human algorithmic innovation.

Put it next to Wednesday's news. OpenAI released 722 mathematics manuscripts generated by its agents and three of them were withdrawn within 36 hours. Epoch is measuring something different and in some ways harder: not whether a model can prove a theorem whose statement is given, but whether it can produce the one idea in a year of post-training research that moved the whole field. On that, the frontier is at 15 percent, and the gap shows up exactly where humans excel, in choosing which experiment to run next rather than in running it.

The design choices are what make the number trustworthy. The target is a single real innovation with a known date, so contamination is checkable. The compute budget is fixed, so this is not a search-harder-and-you-win story. And the task is end to end, including the writeup, so partial credit for a promising idea nobody could reproduce does not count. Anyone claiming automated AI research should be asked for their InnovationEval score, and right now the honest answer is that nobody's agent would have discovered SDPO this January.

Report: https://epoch.ai/publications/innovationeval
Hacker News: https://news.ycombinator.com/item?id=50027257
← Previous
Prime Agent Rewrote Its Harness in Rust: 14x Faster Cold Start, 300,000 Downloads In
Next β†’
Busabase Was #1 on Product Hunt: An Open-Source Database Where Agents Leave Their Work Behind
← Back to all articles

Comments

Loading...
>_