Shapley-Guided Rollout Selection for Reinforcement Learning with Verifiable Rewards
Abstract
Reinforcement learning with verifiable rewards improves language-model reasoning, but reducing training costs through response selection risks losing useful learning signals that reward statistics alone cannot identify. We aim to select a small subset of generated responses that preserves the complete group's learning direction under a fixed update budget. We introduce Shapley Selection Rollouts (SSL), which values each response by its average marginal contribution to recovering the group's direction across different response subsets. SSL measures this contribution using compact token-score representations as proxies for update directions, then passes the highest-valued responses to the existing policy optimizer without changing its objective. Across three language models and seven mathematics benchmarks, including GSM8K, MATH500, and AIME24/25, SSL improves mean benchmark accuracy by 0.49–0.62 percentage points over the strongest evaluated alternative when selecting 16 of 64 responses per prompt. By valuing responses in relation to one another, SSL provides a practical approach to preserving useful learning signals within a limited policy-update budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.