acceptodds
Under review as a conference paper at ICLR 2027

Look Back to Move Forward: Reusing RL-Induced Policy Shifts for Policy Improvement

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves language model reasoning, but obtaining these gains often requires costly trajectory sampling under sparse outcome feedback. This raises the question: can the policy shift learned during an RL run be reused to further improve the same model? We propose Distillation with Implicit Reward (DIR), a framework that turns RL-induced policy shifts into reusable supervision for further policy improvement. Under idealized KL-regularized policy optimization, the log-policy ratio before and after RL is equivalent to the inducing reward up to scale and an additive constant. Motivated by this correspondence, DIR extracts a token-level implicit reward from a fixed checkpoint pair and reuses it in a second regularized optimization anchored at the post-RL policy, yielding an extrapolated distillation target. DIR then distills this target on-policy, providing dense token-level supervision on states visited by the evolving student without a separately trained reward model or an external teacher. Experiments on mathematical reasoning and agentic search demonstrate two complementary benefits. DIR improves performance beyond RLVR. When initialized from intermediate RL checkpoints, it also approaches later-stage RL performance at substantially lower training cost. These results suggest that RL-induced policy shifts can serve as reusable supervision, supporting both continued performance improvement and more compute-efficient post-training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.