October 10, 2026ResearchRLAgents

The Best Paper This Week Finally Answers a Question We've Chased for Months: When Do You Fix the Harness, and When Do You Train the Weights?

For weeks the running argument on this site has been harness versus model. Does an agent get better because you rebuilt the scaffolding around a frozen model, or because the model itself learned something? A paper submitted to arXiv on October 8, "Harness Evolution Hits a Ceiling: When Weight Training Should Begin," is the first one to turn that into a decision rule instead of a vibe. The rule is almost embarrassingly clean: label every failed trajectory by the first thing that went wrong, then separate process failures from content failures. A process failure is a blocked call, a loop, an exhausted step budget, the agent tripping over the plumbing. A content failure is the agent delivering a bad plan or a wrong answer. Process failures are harness work. Content failures are weight work. That is the whole thesis, and the experiments back it hard.

The numbers tell a two-act story. Act one: a self-evolving harness around a frozen Qwen3.5-4B lifts held-out performance on a planning benchmark from 0.16 to 0.30, and the 9B from 0.32 to 0.44, with the 4B's delivery rate climbing from 55 to 90 percent. Evolving the scaffolding works, right up until it doesn't. Act two: train a small LoRA adapter on the trajectories that evolved harness produced, and you get another +0.13 under the original harness for both model sizes. On the 9B the adapter alone matches the entire harness-evolution line, and it cuts content failures from about a quarter of trajectories to roughly 5 percent. A placebo adapter trained on answer-shuffled data falls below the base model, so the gain is real signal, not just more fine-tuning.

The transfer result is the one that makes the rule operational. On WebArena-Lite, 117 unseen tasks, the harness gain lives almost entirely in what the model gets to see, and the adapter adds nothing. Different task, different bottleneck, different fix, and the failure-composition label told you which in advance. Eight models across six families. This is the first time the choice between building scaffolding and spending GPU hours has a cheap diagnostic in front of it: look at how your agent fails, and the shape of the failures points at the fix.

Why this lands now: the same arXiv batch had HarnessSQL taking a text-to-SQL model from 15.5 to 45.2 percent purely by training inside the deployment harness, and Harness Compilation adding up to 24 points to small vision-language students by revising their harness offline. The field is converging on the idea that the harness is part of the training distribution, not a wrapper around it. This paper is the one that tells you when to stop decorating the wrapper and change the thing inside.

Paper: https://arxiv.org/abs/2610.11655
← Previous
Your Coding Agent Can Refactor a Monorepo But Can't Point at a Button. This Gives It a Finger.
Next β†’
A Self-Improving Agent With a Silently Broken Grader Looks Exactly Like One That Works. This Paper Proves It.
← Back to all articles

Comments

Loading...
>_