acceptodds
Under review as a conference paper at ICLR 2027

Decoupling Exploration from Optimization in RLVR

Abstract

Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can discover new ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, degradations are difficult to recover from. We introduce Exploration-Distillation (ExpDis), a framework that decouples exploration from optimization. ExpDis trains an policy with a novelty bonus in the reward, filters its trajectories for correctness and quality, and distills them into a separate policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. Further, parallel explorers can evolve independently to produce diverse distillation targets for the student. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@ scaling, indicating that ExpDis produces models that generate more diverse correct solutions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.