Codex Reset
GitHub

QQ·微信群

CODEX / SIGNAL STUDIO

Codex の動きを、ひと目で。

UT Austin study finds token-cutting context compression can slow coding agents

elvis

A UT Austin study of context compression in coding agents ran nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench. It varied three decisions separately: how context is compressed, when compression triggers, and how much is removed. elvis reports that on Terminal-Bench with Qwen, policies using about a third of the tokens can take 20% to 80% longer than keeping full context.

Step-triggered policies cut the most tokens per step but need 10% to 27% more model calls. Threshold-triggered policies cut tokens by 22% to 55%, with call counts close to full context. Results also differ by model: elvis says a policy that works well for Qwen drops Devstral to 38.7% and makes it slower, so latency and cost should be measured for each model before choosing one. The paper is [Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents](https://arxiv.org/abs/2609.32961).