acceptodds
Under review as a conference paper at ICLR 2027

Single-Transformer Latent Diffusion

Abstract

Latent generative models typically rely on three separate networks: an encoder that tokenizes pixels, a generator that models the latent distribution, and a decoder that maps latents back to pixels. This contrasts with multimodal transformers, where heterogeneous modalities are processed by a single backbone in one unified framework. Recent work narrows the gap by casting tokenization and generation as one latent-inference problem solved by a single generative encoder, but the decoder remains a separate network. We investigate whether this last component can be folded in too, and what it takes for one diffusion transformer to encode, denoise, and decode. We propose a simple formulation that casts all three modes as position-aligned sequence maps over one fully shared backbone; the modes differ only in lightweight input and output heads and in their conditioning, namely a mode token, the timestep, and the class label. On ImageNet-1K at matched parameters, our model, UNIFIED, improves gFID over both a separate tokenizer–generator pipeline and the two-network baseline, with the largest gains at the two smaller scales (2.79 vs. 3.04 at M, 2.44 vs. 2.77 at M) and a smaller gain at M (2.16 vs. 2.23). The results show that a single transformer can encode, denoise, and decode, and that full unification is not merely feasible but can be beneficial: the standalone decoder is a less parameter-efficient use of capacity than the shared backbone, and folding it in can improve generation quality at matched parameters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.