Rethinking diversity in RLVR: from policy entropy to semantic coverage
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for improving the reasoning capabilities of Large Language Models. However, we find that successful responses become increasingly concentrated in fewer semantic reasoning modes as training proceeds. Under a finite rollout budget, this concentration causes on-policy sampling to repeatedly revisit dominant modes, reducing coverage of alternative successful regions and limiting further policy improvement. Theoretically, we attribute this depressing phenomenon to the finite-sample policy gradient estimation, which theoretically leads the policy model collapse on a single response. Motivated by our analysis, to preserve the diverse response modes, we propose SCARS, a framework that organizes discovered responses and intermediate reasoning states in a hierarchical semantic tree and redirects rollout budget toward under-covered successful regions through prefix-conditioned generation. Experiments across multiple reasoning benchmarks show that our method improves overall reasoning performance while maintaining substantially higher successful-response diversity than other baseline methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.