Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
Abstract
Latent generative models typically train a VAE for reconstruction and then fit a generative model on the frozen latent space. Since reconstruction-optimized latents are not necessarily generation-friendly, jointly training both models is an appealing alternative. However, direct end-to-end training remains challenging, as it is prone to latent collapse and faces a generation–reconstruction conflict. We revisit this problem and identify two key insights. First, the entropy term in the Kullback-Leibler divergence (KL) objective is essential for preventing collapse. Second, reconstruction and generation exhibit asymmetric learning dynamics: reconstruction is fast and strongly supervised, whereas generation is slower and harder to optimize. Based on these observations, we introduce an entropy-preserving end-to-end objective for stable joint training and propose GenFirst, a simple generation-before-reconstruction strategy that first shapes a generation-friendly latent space and then refines reconstruction. GenFirst generalizes across exact-likelihood continuous autoregressive priors and SiT-based flow-matching priors. With our end-to-end objective and GenFirst, SiT achieves a gFID of with CFG and without CFG on ImageNet-, while MMDiT reaches a GenEval score of on text-to-image generation. We further extend the framework to shared visual latents for generation, representation learning, and reconstruction, as well as continuous text latents for unified text–image generation, demonstrating its applicability across generative priors, objectives and modalities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.