acceptodds
Under review as a conference paper at ICLR 2027

Beyond Prompt Selection: Learning to Allocate Rollouts for Efficient GRPO Training

Abstract

GRPO, a widely used RL method for reasoning models, allocates the same number of rollouts to every prompt, despite differences in the value of additional rollouts. This is particularly wasteful for degenerated prompts, for which sampled responses receive identical rewards and therefore provide no policy gradient. We introduce AdaRoll, an adaptive rollout allocation framework that dynamically concentrates computation on prompts where additional rollouts are most likely to be useful. In the first training epoch, AdaRoll performs pilot rollout initialization, using one rollout per prompt and its entropy and reward to predict prompt degeneracy and the distribution of future correct and incorrect rollouts. It then progressively allocates additional rollouts over multiple rounds according to a cost-benefit tradeoff between their predicted learning value and generation cost. In subsequent epochs, the predictor continuously adapts to the evolving policy using historical and newly generated rollout information, without requiring additional pilot rollouts. Across model scales and mathematical reasoning tasks, AdaRoll's predictions are 31% more accurate than the baselines, saving up to 47% of computation and improving the LLM's final performance by up to 2.5%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.