acceptodds
Under review as a conference paper at ICLR 2027

From Selection to Switching in RLVR

Abstract

Data selection can improve training efficiency and performance, but its computational costs and effectiveness must be analyzed before it is applied. We compare uniform sampling, on-policy gradient alignment, and cached success-rate (SR) selection in reinforcement learning with verifiable rewards (RLVR). Matched-update MATH experiments under GRPO show selection gains over uniform sampling, with on-policy alignment stronger initially on average and cached SR stronger at later checkpoints. We then examine whether the selection strategy can be changed online during training, from alignment to SR, through executed switching runs and retrospective diagnostics. On the evaluated MATH/GRPO trajectories, switching achieves higher final reward at matched updates than continued alignment and SR started immediately after a shared prefix. Fixed-time switching experiments provide complementary evidence. RLOO controls show objective-dependent selector ordering, while MBPP extends reward and cost comparisons to code generation. With an existing reward cache, incremental selection-and-training costs are lower than continued alignment but higher than early SR. These results support changing selection strategies during RLVR while accounting for selection costs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.