CMU paper trains a 4B proposer to edit agent harnesses without changing the solver
A CMU paper, [Harness Learning Enables Generalizable Test-Time Adaptation](https://arxiv.org/abs/2609.35738), trains a proposer model with reinforcement learning to improve an agent by editing its harness code rather than changing the solver model’s weights.
The proposer reads a task, the current harness, and an execution report, then writes a code edit. Its reward is the score of the revised harness, and the solver model stays fixed.
A trained 4B proposer beats its 35B teacher at single-step revision on Reasoning Gym, including task families it never saw in training. A proposer trained on HotpotQA keeps improving harnesses on MuSiQue and 2WikiMultihopQA.