Look Back to Move Forward: Reusing RL-Induced Policy Shifts for Policy Improvement
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves language model reasoning, but obtaining these gains often requires costly trajectory sampling under sparse outcome feedback. This raises the question: can the policy shift learned during an RL run be reused to further improve the same model? We propose Distillation with Implicit Reward (DIR), a framework that turns RL-induced policy shifts into reusable supervision for further policy improvement. Under idealized KL-regularized policy optimization, the log-policy ratio before and after RL is equivalent to the inducing reward up to scale and an additive constant. Motivated by this correspondence, DIR extracts a token-level implicit reward from a fixed checkpoint pair and reuses it in a second regularized optimization anchored at the post-RL policy, yielding an extrapolated distillation target. DIR then distills this target on-policy, providing dense token-level supervision on states visited by the evolving student without a separately trained reward model or an external teacher. Experiments on mathematical reasoning and agentic search demonstrate two complementary benefits. DIR improves performance beyond RLVR. When initialized from intermediate RL checkpoints, it also approaches later-stage RL performance at substantially lower training cost. These results suggest that RL-induced policy shifts can serve as reusable supervision, supporting both continued performance improvement and more compute-efficient post-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.