CMU paper trains a 4B proposer to edit agent harnesses without changing the solver
The harness-learning paper uses reinforcement learning so a proposer revises harness code from a task, the current harness, and an execution report. A trained 4B proposer beats its 35B teacher on single-step Reasoning Gym revisions, and HotpotQA training transfers to MuSiQue and 2WikiMultihopQA.