ActiveSaddler adapts scenarios while optimizing agent harnesses
ActiveSaddler groups recurring failures into failure patterns and treats each pattern as an arm in a non-stationary bandit. It tracks how much the harness is still learning from each pattern and splits the budget between revisiting known weaknesses and finding new ones.
DAIR.AI says current harness optimizers change how the harness is updated but keep the training scenarios fixed, so feedback continues to come from tasks that stop being informative as the harness improves. ActiveSaddler adapts those scenarios as well. With the same optimizer, test Pass@1 improves by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 compared with a fixed scenario order. The paper is [ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization](https://academy.dair.ai/papers/activesaddler-automated-curriculum-learning-for-agent-harness-optimization-2610.00906).