acceptodds
Under review as a conference paper at ICLR 2027

Divergence controls entropy in distillation

Abstract

Distillation has become a core primitive of large language model training, yet its properties are not yet well understood. We take an entropic perspective on distillation, studying how the entropy of the student depends on the data and the choice of divergence function that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised fine-tuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the task gets too hard, and in between entropy smoothly changes early in training. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.