GuaVAE: Guidance-Aligned End-to-End VAE Training for Latent Diffusion
Abstract
Recent studies have demonstrated that pretrained visual encoders such as DINO can help to improve latent spaces for diffusion modeling. However, these approaches depend on an external representation model. We investigate whether diffusion transformers (DiTs) themselves can provide supervision for learning better latent spaces through VAE tuning. Prior work has shown that classifier-free guidance (CFG) can enhance internal DiT features, making them effective targets for representation alignment during training. We introduce GuaVAE, an end-to-end VAE tuning method that extends this objective to the VAE encoder. To encourage the VAE to learn a more generation-friendly latent space, we align a trainable DiT's features with CFG-enhanced features obtained at low noise levels from a frozen pretrained DiT. We backpropagate the alignment loss through the trainable DiT to the VAE encoder. Experiments show that the tuned VAE improves generation quality across DiT backbone sizes, achieving performance competitive with DINO-supervised VAEs. We also observe better generation quality when using VAEs tuned with larger DiTs, supporting diffusion representations as a scalable source of supervision. A DiT trained with the tuned VAE and CFG-enhanced feature alignment achieves a gFID of on ImageNet without CFG, outperforming self-contained representation alignment baselines as well as DINO-guided REPA. With CFG, it achieves a competitive gFID of .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.