Event-Driven Compute Allocation in Fully Asynchronous LLM Training
Abstract
Reinforcement learning (RL) for large language models is computationally intensive, with rollout generation consuming substantial computational resources. Recent fully asynchronous RL systems improve hardware utilization by decoupling rollout generation from model updates, but they leave open a central algorithmic question: once a rollout job finishes and a generation slot becomes available, how should the system allocate the next unit of compute to maximize learning efficiency? This compute-allocation question is particularly important in large-scale RL runs, where the goal is to achieve the best possible performance under a fixed compute budget. We propose EDCA (Event-Driven Compute Allocation), a unified framework for prompt sampling and rollout budget allocation in fully asynchronous RL. Unlike existing methods, which typically study either prompt selection or adaptive rollout allocation in synchronous settings, our framework treats job completion as the fundamental control point of the learning system. At each completion event, the scheduler chooses between two actions: continue, which launches more rollouts for the current prompt, and replace, which retires the current prompt and fills the freed slot with a newly sampled one. Across three model backbones and six mathematical reasoning benchmarks, EDCA improves average accuracy over fixed-rollout GRPO, including a 2.81% point gain on DeepSeek-R1-Distill-Qwen-1.5B. EDCA effective gradient ratio reaches 92.58%, compared with 56.92% for GRPO. When combined with dynamic sampling, EDCA reduces rejected samples by 76.1% while completing more policy updates within the same wall-clock budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.