acceptodds
Under review as a conference paper at ICLR 2027

A First-Step Bias Restores Lost Diversity in Diffusion Model Distillation

Abstract

Distribution-matching distillation enables high-quality image generation in only a few denoising steps. However, distilled student models often map different initial noise samples to similar outputs under the same prompt. To locate where diversity is lost, we compare the teacher and student along the sampling trajectory. The largest measured reduction in the response to different initial noises occurs during the first, high-noise step. When we replace the first step with the teacher's corresponding trajectory and leave the final three student steps unchanged, most of the output diversity returns. This suggests that the first step is a key point for restoring lost diversity. Motivated by this finding, we introduce *ReSeed*, a post-distillation method that adds a seed-response bias adapter only to the first student step and aligns its velocity with the teacher's first-interval secant velocity. All student weights and later denoising steps remain unchanged. We evaluate *ReSeed* on three text-to-image models (SD3-Medium, SD3.5-Medium, and FLUX.1-dev) and one text-to-video model (Wan2.1-T2V-1.3B). For a distilled 4-step FLUX.1-dev model, *ReSeed* requires only additional trainable parameters relative to the base model and is trained on 8,000 prompt–noise pairs, while keeping the distilled student fully frozen and requiring no target images, perceptual losses, or adversarial losses. It improves DINO diversity by while maintaining competitive visual quality. These results show that substantial diversity can be recovered through lightweight, data-efficient tuning of only the first denoising step, without fine-tuning the full student model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.