Train What Is Ready but Not Yet Mastered: Competence-Driven Curriculum Sampling for Reinforcement Learning with Verifiable Rewards
Abstract
Reinforcement learning with verifiable rewards (RLVR) is an effective paradigm for improving mathematical reasoning in large language models, yet the training prompt distribution is commonly determined by random sampling or manually specified curricula. Such schedules ignore that the useful training difficulty changes with the learner’s competence. We introduce a competence‑driven curriculum sampler that dynamically allocates each training batch across ordered difficulty regions according to two complementary signals: cumulative prerequisite readiness, which estimates whether the model has acquired the easier abilities required to approach a region, and mastery gap, which measures how much competence remains to be learned in that region. We maintain an exponentially smoothed success rate for each difficulty bucket, define readiness as the product of competence estimates over all prerequisite buckets, and assign sampling priority proportional to readiness times remaining mastery gap. This yields a closed‑loop curriculum whose progression is driven directly by model competence rather than training step. On DeepMath‑10K, the same scheduler achieves the highest eight‑benchmark mean among the compared sampling baselines at both Qwen3‑1.7B and Qwen3‑4B scales: 0.3992 versus 0.3944 for the strongest 1.7B baseline, and 0.5617 versus 0.5485 for the strongest 4B baseline. Difficulty‑range ablations further show that the best fixed range changes with model scale, while our adaptive scheduler remains best at both scales. These results support a simple principle for RLVR data selection: train what is ready to learn, but not yet mastered.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.