Extra Dimensions, Simpler Flows: Accelerating Flow Matching Training with Pretrained Visual Representations
Abstract
It is an effective strategy to accelerate the training of diffusion and flow-based generative models by jointly modeling image latent features and semantic representations extracted from pretrained vision encoders. However, stronger encoders do not always improve generation quality and can hinder optimization. We show that a semantic representation reduces velocity ambiguity only when it provides velocity-relevant information that is unavailable from the noisy image state, thereby distinguishing otherwise ambiguous velocity targets. Motivated by this insight, we propose Joint Semantic Flow Matching (JsFlow), which jointly transports the image state and semantic variables from multiple pretrained encoders. JsFlow accelerates semantic transport to provide more informative semantic states throughout the trajectory. Joint image–semantic prediction can still introduce gradient interference through the shared backbone. To address this issue, we predict semantic velocities through a separate branch that takes detached backbone features as input. This preserves semantic guidance during transport while preventing semantic-loss gradients from updating the image backbone. On ImageNet , JsFlow with LightningDiT-XL/1 achieves an unguided FID of after only K iterations. It outperforms SiT-XL/2 trained for M iterations (FID ) and LightningDiT-XL/1 trained for M iterations (FID ), requiring only and of their training iterations, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.