What Does Distillation Transfer? Causal Tests with Ground-Truth Teachers
Abstract
Distilled students can reproduce a teacher's answers without using the same computation. We test what transfers through logits, hidden states, and state traces using modular-addition and automaton teachers with measurable internal signatures. In modular addition, a model's wrong-answer probabilities carry the Fourier frequencies it computes with, and changing only those probabilities selects which frequencies a freshly initialized student uses. Exchanging these probabilities between teachers transfers the source teacher's causal frequency signature. Permuting them while preserving entropy and the correct-answer probability selects a different signature, predicted before training from the edited target's frequency content. Removing the predicted frequencies sharply reduces accuracy; retaining them preserves accuracy in the trained network. A student that already computes through one frequency set can later match a new target and still depend on the old set. Raising the loss weight on those same wrong-answer probabilities restores transfer from near-deterministic teachers. Program-generated automaton traces teach state tracking but contain no teacher identity at training lengths. Using 1.5B teachers and 0.5B students, the probabilities on tokens other than the reference token again determine which teacher's output signature the student acquires. Swapping the output projection moves most of the measured gap. Distillation can therefore select a causal internal signature, but matching the target does not guarantee that the student computes through it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.