VeriHarness turns a base model into a verifier for long-horizon agent tasks
A Google paper, VeriHarness, turns the same base model into an agentic verifier with two jobs. One resolves claims where sampled agent rollouts disagree by checking workspace evidence. The other challenges claims that every rollout agrees on and looks for requirements they all missed. The paper says agreement can hide shared errors, while disagreement often points to the correct alternative.
Across five long-horizon benchmarks, the paper reports the best selection scores among the baselines tested. With evidence-backed revision, it adds 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. The authors also release about 26,000 rollouts. The paper is on [arXiv](https://arxiv.org/abs/2610.00972).