Mitigating Exposure Bias in Flow Matching via Reverse Endpoint-Anchored Supervision
Abstract
Flow matching (FM) has emerged as a powerful generative modeling framework with a simple simulation-free training objective. However, its training and inference procedures exhibit an inherent mismatch: training states are analytically constructed along prescribed interpolation paths, whereas inference recursively evolves through states induced by the model's own learned dynamics. This discrepancy leads to exposure bias, limiting generative performance. Although incorporating inference states into training is straightforward, obtaining reliable supervision matched to these states remains challenging, as their exact ground truth endpoints are generally unavailable in advance under nonlinear model dynamics. To address this challenge, we propose Reverse Endpoint-Anchored Supervision (RES), a post-training framework that generates the proxy of inference states through the reverse dynamics of an online model and aligns them to real inference states via forward dynamics bridging. Their known trajectory origins then serve as real endpoint targets, thereby allowing the online model to train directly on inference states with their corresponding reliable endpoint supervision. Experiments on representative FM models in both latent and pixel spaces, including SiT and JiT, demonstrate consistent improvements in generation quality, validating the effectiveness and generality of our approach.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.