Harness-of-Harness: The Harness Gets Its Own Harness
The name alone earns the click. Harness-of-Harness (HoH), on arXiv September 1 from a nine-author team led by Haoyang Yan, wraps existing coding-agent harnesses — they tested Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3 — in an outer loop that plans, codes, tests, and then improves the software again on the next iteration. Not a new agent; a meta-layer that makes the harness you already have run for days.
Numbers: 52.25% average relative gain across three benchmarks, up to 82.86% after three iterations. The demo that sticks: 70+ autonomous iterations produced a fully playable first-person shooter with a storyline, core mechanics, polished visuals, and audio. The design choices are quietly sensible — balance repairing bugs against growing capability, scope work into small verifiable increments, and expose tools and skills progressively instead of prescribing a rigid workflow.
This slots into a thread we've been tracking all quarter: the harness is where the intelligence gains are cheapest right now (see HarnessLens on evolving harnesses without burning eval budget: https://clauday.com/article/79d486a3-3972-4339-b2c8-2eed15a787ca). HoH's contribution is showing the harness itself can be the unit that gets orchestrated — and that the same outer loop lifts three very different harness-model stacks. When wrapping the wrapper yields 52%, the message is uncomfortable: the models have headroom their scaffolding isn't using.
Paper: https://arxiv.org/abs/2609.01481
← Back to all articles
Numbers: 52.25% average relative gain across three benchmarks, up to 82.86% after three iterations. The demo that sticks: 70+ autonomous iterations produced a fully playable first-person shooter with a storyline, core mechanics, polished visuals, and audio. The design choices are quietly sensible — balance repairing bugs against growing capability, scope work into small verifiable increments, and expose tools and skills progressively instead of prescribing a rigid workflow.
This slots into a thread we've been tracking all quarter: the harness is where the intelligence gains are cheapest right now (see HarnessLens on evolving harnesses without burning eval budget: https://clauday.com/article/79d486a3-3972-4339-b2c8-2eed15a787ca). HoH's contribution is showing the harness itself can be the unit that gets orchestrated — and that the same outer loop lifts three very different harness-model stacks. When wrapping the wrapper yields 52%, the message is uncomfortable: the models have headroom their scaffolding isn't using.
Paper: https://arxiv.org/abs/2609.01481
Comments