acceptodds
Under review as a conference paper at ICLR 2027

Understanding Diversity Collapse in RLVR via the Lens of Overtraining

Abstract

Reinforcement learning with verifiable rewards (RLVR) often suffers from diversity collapse: Pass@ improves while high- Pass@ degrades. We study this divergence through the lens of overtraining. We first show that validation Pass@ largely plateaus while training-set success continues to climb and high- coverage continues to decline. We then show that a problem can retain substantial Pass@ headroom after its direct contribution to high- coverage has nearly saturated. A Bayesian analysis quantifies this distinction under few-rollout uncertainty. Early stopping reduces coverage loss. We also find that stopping updates on high-success problems preserves or improves Pass@ relative to the base model while retaining Pass@ gains, even as training continues on the remaining problems. Together, these results identify overtraining as a source of diversity collapse in RLVR: continued optimization outlasts its coverage benefit, and restricting its period or scope reduces the resulting coverage loss.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.