acceptodds
Under review as a conference paper at ICLR 2027

Future Discrepancy Knowledge Distillation for Diffusion Model Compression

Abstract

Diffusion model capacity distillation trains a compact Student to imitate a high-capacity Teacher, reducing the computation required for each network evaluation during generation. Classic knowledge distillation (KD) typically matches their current predictions at states independently constructed by noising real data. During generation, however, the current state is produced recursively through preceding model predictions and state updates. Different Teacher and Student sampling histories can therefore yield different states for local supervision. Under a common local objective, we compare Teacher- and Student-rollout states and find that their relative performance depends on the Teacher–Student configuration. Even with the training-state source fixed, local prediction matching measures only the immediate discrepancy without accounting for its evolution during subsequent sampling. Local errors of similar magnitude can produce different future state deviations under shared reference dynamics. To address this gap, we introduce Future Discrepancy (FD). Starting from a common state, the Teacher and Student each perform one state update. We then propagate the two next states with the same frozen Teacher and use their discrepancy at a selected future position to supervise the current Student state update. FD applies to either Teacher- or Student-rollout states. We use Temporal Sampling to reduce Teacher continuation queries during training and combine FD with complementary classic KD on independently noised data states to form Future-Discrepancy Knowledge Distillation (\method). Experiments across Student construction strategies, Teacher–Student configurations, and image generation tasks demonstrate the gains of FD over local supervision using the same state source and the effectiveness of the complete FD-KD method.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.