VeriHarness turns a base model into a verifier for long-horizon agent tasks
A Google paper, VeriHarness, uses the same base model to check both disputed and unanimous agent claims. With evidence-backed revision, it reports gains of 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8 over a single rollout.