RL Post-Training Infrastructure
5.29.1Actor, Rollout, Reward, Critic, and Reference Model#
5.29.2Combining the Training Engine and the Inference Engine#
5.29.3Colocated, Disaggregated, and Hybrid Resource Layouts#
5.29.4Synchronous/Asynchronous RL Pipelines#
5.29.5Partial Rollout, Dynamic Sampling, and Stragglers#
5.29.6Multi-Turn Rollout, Environment Interaction, and Tool Execution#
5.29.7Online Weight Update and Weight Broadcast#