dREPA: Difficulty-Aware Representation Alignment Accelerates Diffusion Training
Abstract
Representation alignment (REPA) accelerates diffusion transformer training by aligning intermediate features with a frozen encoder, but it averages both the denoising and the alignment loss uniformly over spatial tokens, although denoising difficulty concentrates on boundaries and fine structure. We show that the per-token alignment loss, which REPA already computes and discards, is a free spatial difficulty signal: tokens that are hard to align are consistently hard to denoise. Turning a difficulty signal into useful weights is not automatic, however. Reweighting the denoising loss by its own per-token residual hurts, and a matched comparison of signals through one pipeline points to two properties a useful signal needs: it should rank tokens consistently across noise draws, and it should stay current with the model. The online residual lacks the first, a frozen signal lacks the second, and the alignment loss has both at no extra cost. We propose dREPA, which converts the alignment loss into bounded, per-sample-normalized weights for the denoising loss in five lines of code with no additional parameters. dREPA is a training-acceleration method: on SiT-XL/2 it reaches at 1M steps the FID that iREPA reaches at 4M, and with guidance tuned independently for each model it improves FID from 2.21 to 2.04. The gain is consistent across model scales, teacher encoders, latent and pixel-space backbones, resolutions, and class-conditional and text-to-image generation, and on SiT-XL/2 it exceeds the run-to-run variation measured in a matched-seed study. Code: https://anonymous.4open.science/r/dREPA-17C2/README.md
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.