Representation Geometry in Diffusion Transformers: What Matters for Generation Quality?
Abstract
Diffusion and flow-matching models are the dominant paradigm for image generation, but are notoriously expensive to train. A recent line of work accelerates training and improves generative performance by shaping the internal representations of these models through auxiliary objectives. However, it is unclear which properties actually matter and how to reliably identify them. To address this, we measure eight properties of diffusion transformer (SiT-XL/2) representations and test whether they track generative performance in three complementary settings: a single training trajectory, 16 REPA and iREPA models that differ only in the vision encoder they are aligned to, and models trained with different representation guidance objectives. Several properties correlate strongly with generative performance, measured by Fréchet Inception Distance (FID), along the training trajectory, yet several of these properties do not separate better from worse models once training has converged. This shows that correlation along a trajectory alone is insufficient evidence that a property matters for generation. We confirm this with interventional experiments, in which turning such properties into auxiliary training objectives yields no or only marginal gains. We find that properties that remain predictive throughout our experiments correspond largely to those that are explicitly or implicitly optimized by recent successful methods. These results offer principled guidance for the design of new representation-shaping objectives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.