The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
Abstract
On-policy distillation (OPD) improves student models by aligning the token distributions induced by the student and teacher logits on student-generated trajectories, and has demonstrated substantial empirical gains. Existing generalized OPD methods enable the student to surpass the teacher through reward extrapolation in the output space. However, the language-model head constitutes an information bottleneck: it discards information encoded in the teacher's intermediate hidden states and propagates sampling noise from token-level log-probability ratios into the amplified learning signal, making output-space extrapolation unstable.We observe that reinforcement learning induces a distributional shift from the base model to the RL-trained model, and that the direction of this shift can be measured throughout the network. Motivated by this observation, we propose **RIDE** (**R**L-**I**nduced **D**irection **E**xtrapolation), which extrapolates RL-induced changes directly in the model's internal representation space.At every layer and token position, RIDE computes an RL-induced representation residual and trains the student through layerwise regression toward targets displaced along this residual. Conditioned on a sampled trajectory, we show that regression toward these targets is equivalent to maximizing a linear directional reward defined by the residual, subject to a quadratic penalty centered at the teacher representation. This interpretation makes explicit how the objective encourages the student to move along the direction of the changes induced by reinforcement learning while penalizing excessive deviation from the teacher.Experiments demonstrate that RIDE substantially outperforms OPRD and matches or even surpasses the teacher model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.