Codex Reset
GitHub

QQ·微信群

CODEX / SIGNAL STUDIO

Codex의 흐름을 한눈에.

#post-training

Meta Superintelligence Labs paper describes a Sharpening Tax in RL post-training

Across 42 base and post-trained pairs, base models with a light harness often solve more agentic tasks at large sample counts than their RL post-trained versions, which still win on pass@1. The authors’ PTGS method sets each prompt’s sampling temperature from estimated difficulty during RL, paying a smaller tax and also raising pass@1.