Google's SHIFT Builds a Different Agent Harness for Every Query, Without Running Any
The harness-versus-model thread has produced a week of evidence that the scaffold decides the score. The natural next question is whether you can pick the scaffold per task. Two papers on the HuggingFace board this week say yes, and disagree about how.
SHIFT, from Google, is the per-query version. A harness is the set of roles, instructions, tools and communication structure wrapped around a model, and the right one depends on the query. Searching alternatives at inference time means running them, which is expensive. SHIFT moves execution out of the loop. A local LLM architect learns a policy over harness-building actions, plus a value function trained on measured executions that predicts a utility balancing accuracy against cost. For each query, Monte Carlo tree search uses those predictions to assemble a harness, and only the chosen one runs. Across 9,193 tasks in six benchmarks with a Gemini 3.5 Flash executor, SHIFT reaches about 80% mean accuracy, beating 17 baselines spanning prompting, prompt optimization and workflow search, and the strongest baseline by 7.2 points. A cheaper mode still beats every baseline while using 32% fewer execution tokens. Choosing structure, instructions and tools jointly beats choosing any one of them by up to 9.1 points.
PluginRSI, from USTC, is the accumulation version. Instead of evolving whole harness programs, where mechanisms are tangled and cannot be reused, it represents a harness as a composition of atomized plugins, improves plugins independently, keeps them in a shared library, and recombines them each iteration. It beats existing harness optimization on software engineering, command-line and QA tasks, the evolved harnesses keep their edge when moved to other solver models without re-optimization, and the plugin library makes the next optimization start faster and converge higher on unseen tasks.
Together they mark a shift in what "the harness" is. It stops being a file a human wrote once and becomes a search space with a learned prior. The number to carry from SHIFT is the one that matters for cost: a predicted harness beats an executed search while spending a third fewer tokens.
Links: arxiv.org/abs/2610.04137 and arxiv.org/abs/2609.32423
← Back to all articles
SHIFT, from Google, is the per-query version. A harness is the set of roles, instructions, tools and communication structure wrapped around a model, and the right one depends on the query. Searching alternatives at inference time means running them, which is expensive. SHIFT moves execution out of the loop. A local LLM architect learns a policy over harness-building actions, plus a value function trained on measured executions that predicts a utility balancing accuracy against cost. For each query, Monte Carlo tree search uses those predictions to assemble a harness, and only the chosen one runs. Across 9,193 tasks in six benchmarks with a Gemini 3.5 Flash executor, SHIFT reaches about 80% mean accuracy, beating 17 baselines spanning prompting, prompt optimization and workflow search, and the strongest baseline by 7.2 points. A cheaper mode still beats every baseline while using 32% fewer execution tokens. Choosing structure, instructions and tools jointly beats choosing any one of them by up to 9.1 points.
PluginRSI, from USTC, is the accumulation version. Instead of evolving whole harness programs, where mechanisms are tangled and cannot be reused, it represents a harness as a composition of atomized plugins, improves plugins independently, keeps them in a shared library, and recombines them each iteration. It beats existing harness optimization on software engineering, command-line and QA tasks, the evolved harnesses keep their edge when moved to other solver models without re-optimization, and the plugin library makes the next optimization start faster and converge higher on unseen tasks.
Together they mark a shift in what "the harness" is. It stops being a file a human wrote once and becomes a search space with a learned prior. The number to carry from SHIFT is the one that matters for cost: a predicted harness beats an executed search while spending a third fewer tokens.
Links: arxiv.org/abs/2610.04137 and arxiv.org/abs/2609.32423
Comments