Knowledge Distillation As Conditional Generation
Abstract
Knowledge distillation (KD) typically treats the teacher representation as a deterministic target, overlooking the inherent uncertainty of knowledge transfer. In this paper, we reformulate KD as a conditional generative problem and propose Conditional Generation for Knowledge Distillation (CondKD), which models the conditional distribution of teacher representations given the student representation. However, directly applying conditional generation to KD faces two challenges: the curse of high-dimensional optimization and the lack of semantic supervision from labels. To address these issues, we introduce a Split Tokenization (SplitTok) strategy, achieving stable and effective unsupervised KD. Additionally, we develop the Distribution Contraction technique to integrate label supervision into the conditional reconstruction process. Our theoretical analysis shows that Distribution Contraction provides a stable gradient surrogate for multi-task learning, enabling efficient supervised training without explicit classification loss on multi-step sampling image representations. With such a simplified and unified objective, our training eliminates heavy hyperparameter tuning and effectively scales to strong training schedules. Experiments demonstrate that CondKD improves over the KL baseline by 16.29% under unsupervised KD on CC3M. With label supervision, our method achieves 82.36% top-1 accuracy with ResNet-50 on ImageNet in 600 epochs without extra data, establishing a new state-of-the-art.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.