Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, yet aggregate accuracy gains can conceal a hidden cost: previously solved problems may regress as training proceeds. We characterize this phenomenon as correct-set turnover, capturing the coupled dynamics of solution acquisition and regression, and make retention an explicit training objective alongside acquisition. In group-relative RLVR, mastered prompts provide no reward gradient when all rollouts are correct, yet can still drift under updates driven by other prompts. This observation motivates a timely-review principle: previously mastered prompts should be revisited before such regression accumulates. We propose ReMind, a retention-aware review mechanism that tracks mastered prompts, periodically revisits them with fresh rollouts, and requeues those that regress. Through pre-rollout batch replacement, ReMind preserves the per-step rollout budget. Across 20 image-text, video, and text-only reasoning benchmarks, ReMind improves average accuracy by 3.81 and 3.44 percentage points over GRPO and DAPO, respectively, while consistently improving retention across modalities and algorithms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.