MindPRO: Accelerating LLM Reinforcement Learning via Robust Partial-Rollout Sampling
Abstract
Reinforcement learning with verifiable rewards (RLVR) typically waits for all responses in a rollout batch before each policy update, allowing a small number of long responses to delay the entire synchronous training step. We introduce MindPRO, a prompt-preserving partial-rollout scheduler that mitigates this long-tail barrier without continuously overlapping rollout generation and policy optimization. Building on quota-based partial-rollout scheduling, MindPRO over-provisions prompt groups, advances once a target quota becomes trainable, and manages never-launched prompts and reusable interrupted trajectories through two checkpointable buffers. Prompt preservation here refers to maintaining prompt identity and eligible partial work across cutoff and recovery, while all selected, terminal, fresh, and partial outcomes remain explicitly accounted for. Rollout collection and optimization remain phase-serialized, while optional stability safeguards mitigate the localized policy mismatch and completion-order bias introduced by prefix reuse. Experiments on mathematical reasoning tasks show verified mean-step reductions of 39.8%–48.6% across three MindRLHF workloads, while final-checkpoint evaluations generally retain or improve downstream reasoning quality. We will publicly release our code to facilitate further development of the MindSpore ecosystem.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.