NVIDIA Mid-Harness raises TerminalBench-Lite Pass@1 from 50.0% to 68.0%
Mid-Harness leaves the generator and harness unchanged and works between them. It samples several candidate shell commands and verifies them before one is run. The paper finds that spending more compute on the verifier is more useful than drawing extra samples.
With GPT-5.6 Sol choosing among eight sampled actions, TerminalBench-Lite Pass@1 rises from 50.0% to 68.0%. A weak verifier gains almost nothing from additional samples. When the small TMAX-9B model checks its own candidates, pairwise comparison works best, and distilling the strong verifier into TMAX-9B helps further. Combining action sampling with trajectory sampling reaches higher success at a lower estimated token cost than sampling full trajectories alone. The paper is [Mid-Harness](https://academy.dair.ai/papers/mid-harness-scaling-actions-between-model-and-harness-for-terminal-agents-2609.39982).