What and When to Align: Hierarchy-Decoupled Representation Alignment for Pixel-Space Diffusion
Abstract
Representation alignment accelerates the training of diffusion models by leveraging pretrained visual representations. However, existing methods treat these pretrained visual representations as uniform alignment targets, overlooking differences in semantic granularity and their roles in diffusion training. We investigate these differences in pixel-space diffusion models and find that alignment targets with stronger component-level semantics are associated with faster early improvements in generation quality, whereas category-level semantics provide complementary object-level supervision throughout training. Based on these findings, we propose HierarchY-Decoupled Representation Alignment (HYDRA), which decouples supervision into Scaffold and Anchor alignments depending on their semantic roles. HYDRA further introduces Position-wise Feature Centering to mitigate position-dependent biases in pretrained features used as alignment targets, and a timestep-conditioned projector to adapt alignment to representations that vary across diffusion timesteps. HYDRA improves generation performance across model scales on ImageNet with classifier-free guidance. JiT-B/16 reaches baseline-level FID with 3 fewer training epochs and further achieves 3.15, while JiT-H/16 improves the final FID from 1.86 to 1.74. Our results highlight that effective representation alignment requires considering not only what representations to align, but also when they should guide generative learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.