acceptodds
Under review as a conference paper at ICLR 2027

SpinningRL: Boosting RLVR Efficiency with Theoretical Reward Gain Modeling and Adaptive UCB Sampling

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective in enhancing the reasoning capabilities of Large Language Models (LLMs), yet it remains painfully inefficient. Uniform sampling wastes computation on non-informative prompts and aggravates the zero-advantage problem in GRPO-like methods. Prior schedulers often select medium-success-rate prompts via learnability-style metrics (e.g., or ), which optimize for **gradient magnitude** rather than **policy improvement**. We instead propose a more direct metric: the KL-constrained **potential reward gain** that is estimable from existing rollouts, and introduce SpinningRL, a non-stationary bandit framework that continuously prioritizes prompts with the highest expected gain under the current policy. Concentrating sampling in this subspace improves sample efficiency and reduces baseline variance, mitigating zero-advantage issues. Experiments show consistent gains on mathematics, logical reasoning and coding tasks under the same runtime budget, with 17%-50% fewer rollouts than uniform sampling at matched performance. The code is anonymously available at [this URL](https://anonymous.4open.science/r/SpinningRL-B666/README.md).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.