acceptodds
Under review as a conference paper at ICLR 2027

DeliberateRL: Shaping the Practice Distribution for Compute-Efficient RLVR

Abstract

Reinforcement learning with verifiable rewards (RLVR) typically assigns a fixed rollout budget to every query, even though success probabilities under the policy vary widely. As a result, many response groups are all-correct or all-incorrect: they consume generation compute but yield no group-relative learning signal. We frame allocation as shaping the practice distribution - how training effort is distributed across queries of different policy competence over the course of training. Inspired by deliberate practice, DeliberateRL is a competence-aware offline framework that estimates each query's success rate once, with a verified inference-only profiling pass, and uses this estimate to decide what to practice, how much to practice each query, and when. We instantiate it with sorted-Group Policy Optimization (sGPO), which filters high-success queries, sizes rollout groups toward roughly one expected success, orders training from easier to harder queries, and, following spaced repetition, revisits a small fixed subset of zero-observed-success queries throughout the curriculum. Across mathematical and scientific reasoning tasks with Qwen- and Llama-family models, sGPO matches the accuracy of uniform and online-adaptive baselines while using 2.5x less compute than uniform allocation over a two-epoch horizon, including the one-time profiling cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.