acceptodds
Under review as a conference paper at ICLR 2027

Knowledge Distillation As Conditional Generation

Abstract

Knowledge distillation (KD) typically treats the teacher representation as a deterministic target, overlooking the inherent uncertainty of knowledge transfer. In this paper, we reformulate KD as a conditional generative problem and propose Conditional Generation for Knowledge Distillation (CondKD), which models the conditional distribution of teacher representations given the student representation. However, directly applying conditional generation to KD faces two challenges: the curse of high-dimensional optimization and the lack of semantic supervision from labels. To address these issues, we introduce a Split Tokenization (SplitTok) strategy, achieving stable and effective unsupervised KD. Additionally, we develop the Distribution Contraction technique to integrate label supervision into the conditional reconstruction process. Our theoretical analysis shows that Distribution Contraction provides a stable gradient surrogate for multi-task learning, enabling efficient supervised training without explicit classification loss on multi-step sampling image representations. With such a simplified and unified objective, our training eliminates heavy hyperparameter tuning and effectively scales to strong training schedules. Experiments demonstrate that CondKD improves over the KL baseline by 16.29% under unsupervised KD on CC3M. With label supervision, our method achieves 82.36% top-1 accuracy with ResNet-50 on ImageNet in 600 epochs without extra data, establishing a new state-of-the-art.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.