Study finds keeping agent harness prompts and tools small cuts costs by up to 3x
A [research paper](https://academy.dair.ai/papers/what-does-a-harness-buy-tokens-mostly-2610.04433) evaluated how changing an agent harness affects performance when holding the underlying model constant, testing Claude Code, mini-SWE-agent, and OpenCode across 447 tasks on SWE-bench Verified. Claude Code and mini-SWE-agent finished within 5 points of each other, and switching the harness changed results about as much as rerunning the same harness, with both altering outcomes on 13% of tasks in a 45-task hard subset.
The findings indicate that keeping harness system prompts and tool schemas compact can cut token costs by up to 3x without degrading accuracy. Cost variations were driven by system prompts and tool definitions being resent at every step, multiplying cumulative token usage as agent trajectories grew longer.