Breaking to Preserve: The Geometry of Few-Step Distillation
Abstract
Few-step distillation is commonly viewed as compression, but this perspective overlooks what the student must acquire beyond the base model. Under six-dimensional input perturbations, six directions capture 99% of response variation in the SD1.5 base model and standard LoRA, compared with 26 for LCM-LoRA and 15 for DMD2 at one perturbation scale. Across six backbones, 90.5–99.2% of the energy in teacher–base response changes lies outside the base span; we call this out-of-span component the teacher margin. Exact matching requires recovering this margin. Any base-confined student therefore faces a matching-error floor, and feature spread alone neither recovers the margin direction nor preserves prompt following. Our breaking–fidelity decomposition reveals distinct learning requirements that equal total errors can obscure and motivates diagnostics of spread, reference alignment, and concentration. For a given starting recipe, these diagnostics predict which control is missing, explaining why the preferred complementary intervention reverses across recipes. On FLUX, complete samplewise and covariance students improve alignment with the teacher's added responses over trained distribution-matching (DM) at all 12 measured pairs and improve generation, while mean base-component error falls and complete response error rises. Translating these diagnostics into training objectives, we compare two implementations; covariance optimizes the diagnosed quantities more directly and outperforms its samplewise counterpart on five of six metrics on each backbone. Across five backbones, our one-step SD1.5 student improves all six metrics over official DMD2, and one-step klein beats its four-step counterpart on Patch FID-T.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.