MixFlow2: One-Stage MixFlow Training for Flow Matching
Abstract
MixFlow reduces the training–testing discrepancy by training the velocity prediction network at a training timestep using training noisy data formed by teacher noisy data at and higher-noise teacher noisy data at slowed timesteps. It is different from standard training: training noisy data is only from the teacher noisy data at the training timestep . The original MixFlow approach adopts a two-stage training strategy: standard training first learns accurate teacher-forced velocity prediction, and MixFlow post-training then uses higher-noise teacher forcing to handle the Slow Flow problem in multi-step discrete sampling. Direct application of original implementation to one-stage training, i.e., training from scratch, is not optimal. The reason is that training noisy data sampling is more from higher-noise teachers, and more capable of handling the Slow Flow problem and less capable of learning teacher-forced velocity prediction. We present a one-stage training approach, MixFlow 2, with the introduction of a novel sampling mechanism: the slowed timestep for forming the training noisy data is sampled from a symmetric triangular distribution, where intermediate slow timesteps are sampled more frequently. In addition, we observe that the prediction at high-noise training timesteps is less reliable, as high-noise timestep is less-frequently sampled. This influences guided sampling: guidance at high-noise timesteps is harmful. So we adopt guidance interval to remove the guidance at high-noise timesteps, maintaining the generation quality achieved without guidance. Empirical results show that MixFlow 2 improves the class-conditional and text-to-image generation performance and outperforms original MixFlow.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.