Training-State Shaping for Low-NFE Flow Matching: Gains Across Solvers and Better Couplings for Distillation
Abstract
Flow matching trains on prescribed interpolation paths, while generation follows trajectories of the learned velocity field. We develop training-state shaping (TSS), a post-training approach that improves few-step sampling and the data efficiency of teacher-generated pairs. In pixel-space models that obtain velocity from clean-image prediction, we reshape training inputs using an exponential-moving-average model while retaining clean targets. We extend TSS to native velocity prediction by pairing each shaped path with its exact tangent. On ImageNet, TSS improves low-NFE generation across JiT model scales and solvers. With JiT-H/16 and DPM++2M, it achieves FIDs of and at 10 and 20 NFE, respectively. In latent space, deterministic reparameterization of training time explains most of the gain. With fixed pools of 200k pairs per source and matched student recipes, shaped-teacher data lets reflow and flow-map students reach target FID with fewer updates, equivalently fewer sample presentations at fixed batch size. We measure this advantage across multiple target FIDs and two teacher pairs. The original flow-map extensions reach similar one-call EMA quality at 100k updates. A separate pool-prefix study finds better student quality from fewer TSS-generated pairs than from a larger vanilla pool. An endpoint-FID-matched control shows that endpoint FID alone does not rank downstream data value.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.