Codex Reset
GitHub

QQ·微信群

CODEX / SIGNAL STUDIO

Your Codex, in focus.

#meta-superintelligence-labs

Meta Superintelligence Labs paper describes a Sharpening Tax in RL post-training

Across 42 base and post-trained pairs, base models with a light harness often solve more agentic tasks at large sample counts than their RL post-trained versions, which still win on pass@1. The authors’ PTGS method sets each prompt’s sampling temperature from estimated difficulty during RL, paying a smaller tax and also raising pass@1.