VeriHarness: When Every Rollout Agrees, That's Where to Look for the Bug
The most counterintuitive line in this week's verification papers is from VeriHarness: disagreement between rollouts often exposes the correct answer, while consensus can hide the errors. Majority voting assumes the opposite. The paper is from Caiqi Zhang, Rujun Han, Zifeng Wang, Nigel Collier, Tomas Pfister and colleagues.
The setup is a fixed base model with no reference answers or rubrics at test time, only multiple rollouts of a long-horizon task. VeriHarness takes the same model the generator uses and turns it into an agentic verifier by giving it a workspace, evidence tools and reusable verification skills. Two roles do the work. A disagreement resolver checks competing claims against evidence in the environment. A consensus challenger attacks the claims every rollout shares and hunts for requirements all of them missed.
Across five long-horizon workspace benchmarks and two frontier models, it gets the best selection scores among the baselines tested. With evidence-backed revision it lifts performance 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. The verification skills also self-improve from failure feedback.
Two things make it worth bookmarking. First, the gain comes from the harness around a fixed model, which is another datapoint for the harness-not-the-model thread. Second, the authors release the full pool of about 26,000 rollouts across all five benchmarks and both models, which they say cost over $100,000 to produce. For anyone building a verifier, that dataset is the expensive part done for free.
Link: arxiv.org/abs/2610.00972
← Back to all articles
The setup is a fixed base model with no reference answers or rubrics at test time, only multiple rollouts of a long-horizon task. VeriHarness takes the same model the generator uses and turns it into an agentic verifier by giving it a workspace, evidence tools and reusable verification skills. Two roles do the work. A disagreement resolver checks competing claims against evidence in the environment. A consensus challenger attacks the claims every rollout shares and hunts for requirements all of them missed.
Across five long-horizon workspace benchmarks and two frontier models, it gets the best selection scores among the baselines tested. With evidence-backed revision it lifts performance 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. The verification skills also self-improve from failure feedback.
Two things make it worth bookmarking. First, the gain comes from the harness around a fixed model, which is another datapoint for the harness-not-the-model thread. Second, the authors release the full pool of about 26,000 rollouts across all five benchmarks and both models, which they say cost over $100,000 to produce. For anyone building a verifier, that dataset is the expensive part done for free.
Link: arxiv.org/abs/2610.00972
Comments