What Visual Generators Need from Teachers: Rethinking Representation Alignment
Abstract
Representation alignment accelerates diffusion transformer training by pulling an intermediate block of the model (student) toward features of a frozen pretrained encoder (teacher). However, which teacher layer to align with, and for how long, are still chosen largely by convention, with each alternative requiring a separate training run. We find that alignment helps where the student cannot linearly recover the teacher's features, rather than where the two representations already resemble each other. Since a deep teacher layer is largely predictable from the layer below it, we isolate what each layer adds—its increment—and measure how much of that increment an unaligned student recovers. The student fills the teacher's hierarchy from the bottom up and stalls near the top, a phenomenon we call hierarchy filling: even after 400K training steps, it recovers almost none of the deepest increments. We define the recoverability gap as the unrecovered share of an increment, which can be estimated from a single unaligned checkpoint. Across short runs that each align one teacher layer at one student block, the recoverability gap nearly reproduces the ranking of teacher layers by FID improvement, while CKA, a measure of feature similarity, largely reverses this ranking. Based on these findings, we introduce Representation Alignment and Recoverability Estimation (RARE). Before training, RARE selects the teacher layer with the largest recoverability gap. During training, it tracks each token's remaining distance to that layer as an online counterpart of the gap, weights tokens according to this distance, and phases out the alignment loss once the average distance stops decreasing. With SiT-B/2 on ImageNet 256×256, RARE achieves an FID of 18.02 without classifier-free guidance and 4.46 with guidance, outperforming seven representation-alignment baselines, including REPA, iREPA, and HASTE. RARE also requires 14% fewer GPU-hours than iREPA. Its FID remains lower than iREPA's across model scales, teacher encoders, datasets, and diffusion backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.