Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
Abstract
On-policy distillation (OPD) has emerged as a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher to the student, yet what OPD actually distills into the student's internal representations remains unclear. In this study, we examine this question through the lens of sparse crosscoders, which learn one feature dictionary shared by the student before and after OPD and the teacher. Standard crosscoder analyses, however, identify features specific to one model but cannot tell how a model's use of its features changes, since the crosscoder encodes all models into a single set of feature activations. We therefore propose the swap readout, which reads the feature activations of each student checkpoint on its own and thus measures how training changes the student's use of each feature, even for checkpoints the crosscoder has never seen. Specifically, across three OPD settings, we observe that OPD neither creates features nor passes on the teacher's own, and leaves the firing rates of over of the student's frequently used features within . These results suggest that OPD primarily reweights the features the student already shares with the teacher rather than acquiring new ones. We further examine the SFT warm-up on the teacher's rollouts that commonly precedes OPD and makes it more effective. Rather than adding features, as one might expect, the warm-up reweights the shared ones, partly along OPD's direction, making part of OPD's change in advance, and partly in directions that OPD does not take and that persist through OPD. Imposing this reweighting on the features of a directly distilled student, without changing its weights, brings its accuracy close to that of the warmed-up student, whereas the same change on shuffled features does not. Together, these findings suggest that OPD behaves more like a reweighting of existing features than an acquisition of new ones: the student learns from the teacher how to use the features they already share.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.