Step Ahead, Distill Back: One-Step RL-Guided On-Policy Distillation
Abstract
On-policy distillation (OPD) offers dense supervision for reasoning models, but its use with reward-trained experts commonly follows a train-then-distill workflow. We show that effective guidance can instead come from a teacher only one reward-gradient step ahead of the student. Our method, On-Policy Gradient Distillation (OPGD), rebuilds this teacher from the current actor at every iteration, allowing it to inherit accumulated learning while adapting to fresh verified outcomes. The teacher–actor log ratio isolates the preferences induced by this adaptation and supplies signed token advantages for OPD on the same rollouts. A first-order analysis connects these preferences to interactions between token score gradients, explaining how response-level feedback is redistributed within and across responses. Distillation reweights the outcome gradient through the policy's local geometry while preserving nonnegative first-order alignment. On Qwen3-4B, OPGD improves mean Pass@1 over GRPO by percentage points across four mathematical benchmarks at a common -update budget. Separate training on coding prompts also yields faster learning and higher accuracy across three code-generation benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.