Codex Reset
GitHub

QQ·微信群

CODEX / SIGNAL STUDIO

你的 Codex,盡在掌握。

AutoCompact trains coding agents to decide when to compact context

elvis

AutoCompact trains a coding agent to decide when to compact context, which working state to keep, and how to resume. A judge reviews the base agent’s compaction decisions and replaces flawed ones before they run. The corrected trajectories are used for supervised fine-tuning, and reinforcement learning with task-success rewards then trains coding and compaction together. The paper is on [arXiv](https://arxiv.org/abs/2610.02163).

Pass rates improve by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified. The gain holds even with a 256K window that never overflows, so learned compaction still helps when context space is not the limit. The authors describe this proactive compaction as model-harness co-design: the harness provides the compaction mechanism, while the model learns when to invoke it, what to preserve, and how to continue.