Diversity Collapse in RLVR: Even Good Strategies Die if They Start Rare
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves reasoning accuracy, but it can also narrow the set of strategies a model uses. We ask which strategies survive this process and why. Across a full RLVR trajectory on Countdown and three base/RL model pairs on MATH-500, we find a consistent rich-get-richer effect: strategies that are already common become more dominant after training, while rare strategies are much more likely to disappear. Strategy quality provides weaker and less consistent protection; even reliable strategies can be lost when they begin uncommon. On Countdown, the effective number of strategies per problem falls from about to nearly , while problems whose strategies start at similar frequencies retain substantially more diversity. We further find that finite training samples can amplify this imbalance; once a strategy becomes sufficiently rare, it may stop appearing in sampled rollouts and receive no direct reinforcement. These results show that RLVR can favor what is already common over what is merely effective, causing early frequency differences to compound into the loss of useful reasoning strategies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.