Diffusion as a Multi-Scale Discriminator: Rethinking Noise Scales for Diffusion Model Distillation
Abstract
We propose MuSD (Multi-Scale Discriminator), a simple framework for diffusion distillation that uses a frozen diffusion teacher to guide adversarial training. We train lightweight discriminator heads on the teacher's features at multiple resolutions, without training an auxiliary score network or using a separate pretrained feature extractor. Prior work that uses the teacher as the discriminator biases its noise levels toward high noise to preserve global coherence. We observe that fine-scale structure in teacher features survives mainly at low noise, whereas the coarse focus that features acquire at high noise is also available from coarse readouts at low noise. Building on this, we derive a Beta noise sampler as the maximum-entropy density under a corruption budget, which favors low noise while keeping positive density over the whole noise interval, and we show that both properties are needed. On CIFAR-10, AFHQv2, FFHQ, and ImageNet, MuSD achieves competitive or better one-step FID than the baselines with fewer training images and lower GPU memory use. It applies to both U-Net and Transformer teachers, including EDM2-S and DiT-XL/2, where the one-step student surpasses its 250-step teacher, and extends to two-step text-to-image generation with SD3.5-Medium.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.