Flow-Guided Wasserstein Compensation Distillation
Abstract
Few-step flow distillation constructs training directions from model predictions at student-generated states, but these directions need not realize a prescribed evolution toward a reference distribution. We introduce Flow-Guided Wasserstein Compensation Distillation (FWCD), a framework that combines existing flow guidance with distributional information from unpaired reference samples. At each noise level, FWCD defines the minimum-energy correction to a scaled flow prior that completes a prescribed marginal response. A weak primal–dual formulation learns a scalar potential from unpaired samples; its input gradient supplies the correction for magnitude-controlled student updates. Under suitable function-space assumptions, we characterize the optimal field and its residual decomposition, relate population dual suboptimality to field approximation error, and establish Wasserstein contraction for the corresponding ideal distribution dynamics. Experiments on ImageNet-256 and SANA-1.6B demonstrate improvements over Mean Flow Distillation with four-step inference at matched student and auxiliary update counts. The flow-prior variant reduces ImageNet FID from 4.551 to 4.222 and SANA teacher-reference FID from 27.048 to 24.755. Ablations show task-dependent objective preferences: additive compensation achieves lower mean FID and KID on ImageNet at the default fitting budget, while flow-guided compensation performs better on SANA under a shared cap. FWCD preserves the student architecture and requires no additional model evaluations at inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.