CoreWeave introduces RL Rollouts for live model weight updates in reinforcement learning
CoreWeave has launched RL Rollouts to reduce latency during reinforcement learning training loops. Instead of redeploying and pausing the trainer at each checkpoint, the system loads new weights into a live deployment without interrupting in-flight requests, operating roughly 15 times faster than standard redeployment cycles.
NVIDIA and You.com used RL Rollouts to post-train Nemotron 3.5 Lightning for web search. According to CoreWeave's [breakdown](https://crwv.co/utdYc), the workflow raised BrowseComp accuracy from 36.97% to 45.45% while reducing tool calls by 30.24%.