acceptodds
Under review as a conference paper at ICLR 2027

Where the Prediction Goes Matters in Token-Level Knowledge Distillation

Abstract

Token-level knowledge distillation transfers a teacher’s next-token distribution to a smaller student. Common distillation objectives measure how closely the student matches the teacher’s next-token distribution at each prediction position. However, a smaller student cannot in general reproduce the teacher distribution exactly at every position, raising a question: when the student redistributes probability mass relative to the teacher, should the destination of this mass matter? In particular, should equal amounts of this mass be treated the same when they go to different tokens? We argue that their significance should depend on how the teacher relates those tokens across a population of contexts. To model these relations, we represent each token by a teacher-derived distribution over contexts. At each prediction position, we use the teacher and student next-token distributions to define two mixtures of the teacher-derived context distributions and compare them with the Cauchy–Schwarz divergence. We combine this divergence with a standard probability-matching objective to introduce residual composition distillation (RCD). Across two teacher–student model families and multiple evaluation sets, RCD consistently improves prediction quality and generation diversity, with a trade-off in generation quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.