acceptodds
Under review as a conference paper at ICLR 2027

Align Before Distilling: Teacher-Aligned Repair for Compact One-Step Diffusion Models

Abstract

Pruning and step distillation are typically applied separately: pruning reduces the network size of diffusion models, while step distillation reduces their sampling cost. When combined, however, pruning can leave the student poorly initialized for distillation, which often causes step distillation to fail. In our EDM2-XS experiment, direct SiDA training from a pruned checkpoint produces unusable samples. We introduce Align Before Distilling (ABD), a teacher-aligned repair stage before step distillation: the compact model matches teacher outputs on noisy real-image latents, then uses the repaired weights to initialize one-step training. We evaluate this procedure with SD1.5/SiD-LSG, EDM2-XS/SiDA, and CIFAR-10 EDM/Diff-Instruct. On ImageNet-512, the original EDM2-XS teacher uses 124.713M parameters and 63 network evaluations at FID 3.53. Our 20% target reduction (20.7571% actual) model reaches FID 3.12 with 98.826M parameters and one evaluation; the 30% target reduction (29.4146% actual) model reaches 4.26 with 88.029M parameters. These compact models trade quality for size relative to dense one-step SiDA at FID 2.228. The SD1.5 model reduces parameters by 23.9448% and reaches FID 18.7, while the 36.807M-parameter CIFAR-10 student reaches FID 5.32 with Diff-Instruct.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.