Harness-Aware Distillation trains small model agents to 63.4% success on ALFWorld
Researchers introduced Harness-Aware Distillation, a framework designed for small language model agents operating with an execution harness, according to an [arXiv paper](https://arxiv.org/abs/2610.02858). The method queries the same teacher model with and without harness information, training the student model to prefer actions chosen when harness data is present. A filter discards candidate pairs where the preferred action contradicts harness records, operating without task rewards or success labels.
The researchers observed that simply adding harness data to on-policy distillation raised the student's harness usage on ALFWorld from 65.7% to 73.1%, but left task success flat at 43.1% to 43.5%. With Harness-Aware Distillation, the student model achieved a 63.4% success rate on unseen ALFWorld tasks, outperforming the best baseline at 47.0% and surpassing its 8B teacher. The trained student also escaped 59.7% of stalls, compared to baselines remaining near the untrained student baseline of 46.8%.