acceptodds
Under review as a conference paper at ICLR 2027

Train Less, Learn More: Adaptive Efficient Rollout Optimization for Group-Based Reinforcement Learning

Abstract

Group Relative Policy Optimization (GRPO) assigns a fixed number of rollouts to each query. With binary rewards, all-correct and all-incorrect groups have zero group-relative advantages, while mixed groups may contain more responses than are needed for an effective update. We introduce Adaptive Efficient Rollout Optimization (AERO), which probes queries, reallocates a shared rollout budget to difficult groups, and reduces the responses retained for training. A Bayesian posterior supplies nonzero policy scores for groups that remain homogeneous. On Qwen2.5-family models for mathematical reasoning and single-turn coding, AERO reduces estimated generation-plus-update compute by about 48% and per-step wall-clock time by about 45%, under the same nominal rollout cap, while maintaining comparable Pass@8 and Avg@8. We also test adaptive rollout allocation alone in multi-turn coding, without posterior scores or rejection sampling. With Qwen3.6-27B, the schedule at reaches a 60.8% observed peak pass@1 on a 51-problem SWE-bench Verified split, versus 51.0% for the fixed- GRPO reference, with 2.80 recorded rollout steps per hour. Different update-group targets make this component-level evidence, not a compute-matched evaluation of full AERO.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.