Kavli Affiliate: Wei Gao| Summary:Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill–decode (PD) configuration. Both the choice between […]
Continue.. PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning