acceptodds
Under review as a conference paper at ICLR 2027

Representation Conditioning for Generation

Abstract

Representation-alignment methods effectively accelerate the diffusion transformer training. REPA, for example, uses a pretrained vision encoder to supervise an intermediate layer through an *MLP projection head* which is discarded after training. It aims to import semantics from the vision encoder to the diffusion backbone. While this objective is achieved at *early layers* by back-propagating the alignment loss, there is no guarantee that *later layers*, supervised by the velocity prediction loss alone, can well access such semantics. We propose a simple solution, representation conditioning (RECON), to improve semantics in later layers. During training, we feed the pretrained vision encoder tokens into later layers; during inference, because we no longer have the clean image, we reuse the MLP head (as a proxy to the vision encoder) and project its output tokens to the later layers. Our method adds no parameters and leads to much faster convergence: RECON is faster than SiT and faster than REPA.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.