OPTD: On-Policy Truncated Distillation
Abstract
On-policy representation distillation (OPRD) trains a student language model on its own samples to match its teacher’s hidden states at every layer, giving every layer’s squared error the same weight. Hidden states grow in scale with depth, so this equal weighting sends most of the gradient on the hidden states to the last few layers. Splicing the student’s hidden state into the teacher one layer at a time shows that errors in these layers damage the teacher’s prediction far less than the loss implies. Weighting layers by the measured damage improves accuracy, yet a shuffled placebo that assigns the same weights to random layers improves it even more. Every reweighting scheme we test, this placebo and a uniform one on normalised errors included, starves the last layers of gradient. On-policy truncated distillation (OPTD) keeps only that shared property by dropping the last four layers from OPRD’s sum. With a public teacher trained from its student by reinforcement learning, OPTD improves accuracy over OPRD by 3.97 points on problems from four mathematics and science benchmarks, beats every reweighting scheme, and outscores the teacher on all four.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.