GTA: Global Token Alignment as a Bridge from Autoencoding to Denoising
Abstract
Latent generative models are jointly limited by two factors: the quality of the latent space and the modeling capacity of the denoiser. Although recent representation-enhanced methods improve tokenizers or denoisers, the two stages still lack a shared semantic bridge. Global semantics formed during autoencoding may not be effectively inherited by the denoiser. During denoising, the evolving global semantics may also drift from the spatial latent representations. We introduce GTA, a Global Token Alignment framework that bridges autoencoding and denoising through a shared global representation. GTA encourages the autoencoder to internalize image-level semantics during latent formation and maintains consistency with spatial latents throughout denoising. Experiments on ImageNet 256×256 show that GTA accelerates training and improves generation quality. With SD-VAE initialization, GTA achieves a gFID of 6.42 after 20 epochs. This outperforms REPA-E, which reaches 7.17 after 40 epochs. With VA-VAE initialization, GTA further achieves a gFID of 1.87 after 200 epochs. GTA also produces a more generation-friendly tokenizer. When GTA-VAE is frozen and new generative models are trained from scratch, it consistently outperforms the corresponding various types of VAEs under SiT, REPA, and REG.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.