Single-Transformer Latent Diffusion
Abstract
Latent generative models typically rely on three separate networks: an encoder that tokenizes pixels, a generator that models the latent distribution, and a decoder that maps latents back to pixels. This contrasts with multimodal transformers, where heterogeneous modalities are processed by a single backbone in one unified framework. Recent work narrows the gap by casting tokenization and generation as one latent-inference problem solved by a single generative encoder, but the decoder remains a separate network. We investigate whether this last component can be folded in too, and what it takes for one diffusion transformer to encode, denoise, and decode. We propose a simple formulation that casts all three modes as position-aligned sequence maps over one fully shared backbone; the modes differ only in lightweight input and output heads and in their conditioning, namely a mode token, the timestep, and the class label. On ImageNet-1K at matched parameters, our model, UNIFIED, improves gFID over both a separate tokenizer–generator pipeline and the two-network baseline, with the largest gains at the two smaller scales (2.79 vs. 3.04 at M, 2.44 vs. 2.77 at M) and a smaller gain at M (2.16 vs. 2.23). The results show that a single transformer can encode, denoise, and decode, and that full unification is not merely feasible but can be beneficial: the standalone decoder is a less parameter-efficient use of capacity than the shared backbone, and folding it in can improve generation quality at matched parameters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.