October 2, 2026ResearchCodingAgents

Mid-Harness: Check the Command Before You Run It, +18 Points on Terminal Tasks

Terminal agents have a problem that chat models do not: a bad action changes the world. Install the wrong package and every later step works in a broken environment, even if the model could have proposed the right command on another try. "Mid-Harness" (arXiv 2609.39982, 96 upvotes on Hugging Face) puts compute exactly at that point. It samples several candidate actions, verifies them, and only then forwards one to the harness. The generator and the harness stay unchanged. The new layer sits between them.

The headline: on TerminalBench-Lite, with a TMAX-9B generator and a GPT-5.6 Sol verifier, Pass@1 goes from 50.00% to 68.03% with 8 sampled actions. Same model writing commands, same harness running them. The 18-point gain comes entirely from choosing better among options the small model was already capable of producing.

The finding that matters is about the verifier. With a weak verifier, sampling more actions barely helps. The useful alternatives are there, and nothing can tell them apart. A strong verifier unlocks them. When the same 9B model has to be its own verifier, pairwise comparison works better than the other mechanisms tested. Distilling the strong verifier's judgments into the small model improves it further, without touching the generator.

The cost result is the practical one. Combining action-level and trajectory-level scaling reaches higher success at lower estimated token cost than just generating more full trajectories. Most "test-time compute" for agents today means rerunning the whole episode and picking the best. This says the cheaper unit is the single action, checked before it can do damage.

It fits a pattern from the past few weeks: the gains are coming from the harness, not the model. And it shows exactly where a decision model belongs, as a fast pairwise verifier in front of every irreversible command.

Link: arxiv.org/abs/2609.39982
← Previous
False Frontiers: Self-Evolving Agents Learn to Agree With Themselves, Not the Truth
Next β†’
MILO Evolves Its Own Harness and Tops Terminal-Bench With 26% Fewer Tokens
← Back to all articles

Comments

Loading...
>_