MiTraFlow: Learning Corrective Vector Fields on Model-Induced Trajectories
Abstract
Diffusion and flow-matching models generate images by progressively transforming noise into data samples. However, standard training supervises ideal noising or interpolation states, whereas sampling follows states induced by the model's own predictions. This mismatch creates exposure bias: prediction errors can shift subsequent states away from the training distribution, where recovery is not directly supervised. We introduce MiTraFlow, a flow-matching objective that learns corrective velocity fields on model-induced trajectories. A controlled one-step rollout mixes the predicted and paired target velocities to construct a later state, at which the model learns the velocity needed to reach the paired data endpoint in the remaining time. Crucially, gradients propagate through both the rollout state and its corrective target, allowing the later loss to jointly optimize the correcting prediction and the earlier prediction that produced the state. We further align shallow features at the earlier state with detached deep features at the later state to regularize representations along the same transition. Our analysis shows that the corrective target compensates for the preceding velocity error, and ablations confirm that both controlled error exposure and cross-time gradients are essential. MiTraFlow adds no inference cost and yields larger gains at fewer sampling steps. On ImageNet , it converges approximately faster than REPA and achieves an FID of 1.29 with guidance; the improvements also generalize to pixel and RAE spaces and to large-scale text-to-image generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.