Representation-Augmented Diffusions
Abstract
Diffusion models achieve state-of-the-art sample quality but lack an inherent mechanism to learn semantic representations from data. We propose Representation-Augmented Diffusions (RAD), a new class of diffusion models that natively embeds compact representations within the denoising process. Training maximizes a variational lower bound on the data likelihood, in which a recent class of models based on representation alignment emerge as special cases. Interestingly, our framework reveals a design choice overlooked by prior work: feeding the learned representation back into the denoiser. This simple idea leads to an order of magnitude gains in training convergence speed, data and sampling efficiency over prevalent representation alignment framework while enabling semantic discriminative representations. Furthermore, RAD extends naturally to other modalities like text-to-image synthesis, demonstrating that strong representations can be learned without compromising sample quality. Overall, RAD emerges as a strong alternative for efficiently training diffusion models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.