acceptodds
Under review as a conference paper at ICLR 2027

Understanding the Benefits of Immature Teachers in Knowledge Distillation through Alignment and Regularization

Abstract

Knowledge distillation (KD) trains a small student model on the outputs of a large teacher model.Its effect has mainly been explained from two perspectives: the transfer of inter-class similarities contained in teacher predictions and regularization by soft targets.However, because both effects arise from the same distillation loss, what determines their balance remains unclear.Moreover, reports that immature teachers taken during training can yield better students than mature teachers are difficult to explain from the knowledge-transfer view that more accurate teachers are more beneficial.We show that, in the high-temperature limit, this balance is determined by the teacher logit norm, and we examine the resulting hypothesis experimentally.In this limit, the distillation loss separates into a teacher-independent penalty on the student logit norm and an alignment term whose gradient is proportional to the teacher logit norm.On four image and text classification datasets, the test loss of students distilled from teachers of different maturities followed a nearly common U-shaped curve against the student logit norm, and teacher maturity changed the point that students reached on this curve.Fixing the direction of the mature teacher's logits and replacing only their norm with that of an immature teacher reduced the student test loss below that obtained with the mature teacher, without changing the teacher's predicted classes.At the same teacher logit norm, however, student performance differed with the direction.These results suggest that the balance between knowledge transfer and regularization shifts with the teacher logit norm, and that this balance partly underlies the effectiveness of immature teachers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.