Useful Before It Generates Well: Model-Aligned Sparse Sinkhorn for Flow Retraining
Abstract
Partially trained flow models are usually judged by sample quality, yet they may already carry enough coarse geometric structure to improve later training. We ask whether such early-stage models can be useful before they generate well, by serving as coupling constructors for flow retraining. The available coupling choices fail differently. Independent coupling is simple but high variance. Rectified-Flow-style model-generated coupling reduces variance but inherits teacher and solver drift. Optimal-transport couplings are principled but too expensive at dataset scale. MASS (Model-Aligned Sparse Sinkhorn) uses model-generated endpoints from an early-stage flow only to define candidate neighborhoods in the real dataset, then applies semi-relaxed Sinkhorn reweighting on the resulting sparse graph to produce a real-data-anchored transport plan at cost rather than . The plan is built once before retraining and reused across all training steps, unlike minibatch OT methods that re-solve an OT problem at every batch, and it composes with any flow-matching objective. We characterize sufficient conditions on the neighborhood size under which the MASS empirical target marginal provably improves over the pretrained model's output distribution. On a partially trained CIFAR-10 teacher, the induced-target FID drops from ( NN) to ( MASS), and CIFAR-10 retraining FID improves from (iMF) and (Dec-iMF) to (MASS-ReiMF) at matched compute. On 2D moons, MASS achieves transport cost within of full global OT.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.