MILO Evolves Its Own Harness and Tops Terminal-Bench With 26% Fewer Tokens
If harness design decides how well an agent does, the next step is obvious: let agents design the harness. MILO (arXiv 2609.38349) does this, and it beats every hand-built harness it was compared against. MILO stands for Meta-evolutionary Island Orchestration. It evolves complete agent harnesses and, at the same time, evolves the strategy it uses to search for them.
Three parts make it work. First, a lineage memory over island-based trees that keeps rejected mutations as negative evidence, so the search remembers what failed. Second, a mutator agent on each island that rewrites the whole harness using both the global search history and feedback specific to its parent. Third, an orchestrator that changes the search itself as it runs: grafting lineages, splitting off new species, reassigning mutators and revising the curriculum. Earlier methods tuned one component, a prompt or a skill, or got stuck with one fixed search strategy.
The results with Opus 4.8 are hard to argue with. Starting from the same initial harness, MILO improves resolution by +12.0% on Terminal-Bench 2.1, +28.3% on PaperBench and +10.3% on DeepSWE. The best previous search methods managed +4.5%, +18.3% and 0%. On Terminal-Bench 2.1 it reaches 86.1% ± 2.0%, above the official leaderboard's top entry at 83.8% ± 2.3%, while using 26% fewer tokens than the harness it started from. Better and cheaper at once. It also works with open-weight gpt-oss-120b, and it outperforms eight state-of-the-art harnesses and six search methods.
There is a strange bonus. Pointed at EinsteinArena's open math problems, the same machinery improved the best-known upper bounds for Erdős's minimum-overlap problem and two autocorrelation inequalities. The margins are tiny, in the sixth or seventh decimal place, but they are new records.
Put it next to this week's Mid-Harness and Meta-Skill papers and a pattern is hard to miss. Researchers keep finding double-digit gains without touching model weights. Harness engineering is turning into a search problem, and the searcher is another agent. If your agent product's edge is a hand-tuned harness, expect a machine to tune a better one soon.
Link: arxiv.org/abs/2609.38349
← Back to all articles
Three parts make it work. First, a lineage memory over island-based trees that keeps rejected mutations as negative evidence, so the search remembers what failed. Second, a mutator agent on each island that rewrites the whole harness using both the global search history and feedback specific to its parent. Third, an orchestrator that changes the search itself as it runs: grafting lineages, splitting off new species, reassigning mutators and revising the curriculum. Earlier methods tuned one component, a prompt or a skill, or got stuck with one fixed search strategy.
The results with Opus 4.8 are hard to argue with. Starting from the same initial harness, MILO improves resolution by +12.0% on Terminal-Bench 2.1, +28.3% on PaperBench and +10.3% on DeepSWE. The best previous search methods managed +4.5%, +18.3% and 0%. On Terminal-Bench 2.1 it reaches 86.1% ± 2.0%, above the official leaderboard's top entry at 83.8% ± 2.3%, while using 26% fewer tokens than the harness it started from. Better and cheaper at once. It also works with open-weight gpt-oss-120b, and it outperforms eight state-of-the-art harnesses and six search methods.
There is a strange bonus. Pointed at EinsteinArena's open math problems, the same machinery improved the best-known upper bounds for Erdős's minimum-overlap problem and two autocorrelation inequalities. The margins are tiny, in the sixth or seventh decimal place, but they are new records.
Put it next to this week's Mid-Harness and Meta-Skill papers and a pattern is hard to miss. Researchers keep finding double-digit gains without touching model weights. Harness engineering is turning into a search problem, and the searcher is another agent. If your agent product's edge is a hand-tuned harness, expect a machine to tune a better one soon.
Link: arxiv.org/abs/2609.38349
Comments