Asymmetric Diffusion Transformer: Representation Supervision via Training-Time Queries
Abstract
Pretrained representations are now widely used to train diffusion transformers, and the open question is not whether to use them but what role to give them. Existing methods answer in two ways: representation alignment (REPA) keeps the representation outside the transformer and pulls an internal activation toward it with an explicit loss, while joint diffusion brings it inside but makes it a variable the model must generate alongside the latent. We introduce **Asym-DiT**, a third option: the representation enters the transformer as a training-time query, yet never conditions the generative path, so it is dropped at inference with no added cost. An asymmetric mask lets representation tokens read the latent while the latent never reads them; the representation solves its own denoising task by reading the latent, and its loss reaches the latent only through backpropagation, reshaping the latent hidden states to expose the information the representation needs. The representation's information enters the latent through this query-induced learning rather than an explicit loss: the representation does not tell the latent what to be, only what information to expose, leaving its form free. Because the objective never fixes that form, it does not compete with generation the way an alignment loss does, and it keeps helping where REPA stalls, scaling better with model size and, unlike REPA, giving large gains when fine-tuning an already trained transformer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.