When More Calls Help: Dataset-Free Low-Compute Self-Distillation for MeanFlow
Abstract
Distillation has become an effective way to compress multi-step generative models into fewer steps while retaining much of the teacher’s sample quality. We observe a complementary opportunity in MeanFlow: selected multi-step samplers of an already one-step-capable checkpoint produce better samples than its one-step mode. Together, these observations suggest a simple route to a stronger one-step model: distill the better multi-step mode back into one step, keeping the quality lost in distillation smaller than the quality gained by sampling. We call this recipe **Many-to-One self-distillation (M2O)**. This transfer requires only a small adaptation budget. Using only random noise, class labels, and a frozen self-teacher from the same checkpoint, M2O matches generated distributions in a frozen visual feature space **without accessing training data**. On ImageNet-256, M2O improves iMF-B/2 from 3.35 to 3.00 FID-50K in 0.60 H200-hour. On iMF-L/2, it reaches 1.72 in 0.93 H200-hour, improving on the guidance-calibrated one-step baseline of 1.88 by 0.16 FID. A fixed-guidance control improves from 2.31 to 1.72 at CFG 5.00. Both students preserve single-step, single-branch inference. Our results show that a checkpoint’s multi-step advantage can become a practical source of low-compute, dataset-free improvement for one-step generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.