Codex Reset
GitHub

QQ·微信群

CODEX / SIGNAL STUDIO

Codex の動きを、ひと目で。

Meta Superintelligence Labs paper describes a Sharpening Tax in RL post-training

DAIR.AI

Meta Superintelligence Labs report that base models with a light harness often solve more agentic tasks than their RL post-trained versions when both get enough samples. Post-trained models win on pass@1, but at large K the base models frequently solve tasks the post-trained ones never solve on BFCL v4 multi-turn, ACEBench, and WebShop.

The authors say post-training pushes each task toward always solved or never solved, raising consistency while lowering coverage. They call the lost test-time scalability the Sharpening Tax. Across 42 base and post-trained pairs, it appears in most settings, grows with model size, and can be estimated from a few rollouts.

Their fix, PTGS, sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax and also raises pass@1. The paper is [Sharpening Tax in post-training](https://academy.dair.ai/papers/sharpening-tax-in-post-training-2610.01509).