Purify Before You Align: Improving Representation Alignment for Diffusion Models
Abstract
Representation alignment improves and accelerates diffusion training by encouraging intermediate features to match a pretrained visual encoder. Yet choosing the right encoder remains largely empirical: DINOv2 provides substantial gains, whereas MAE, despite retaining more reconstructive information, offers little benefit. We introduce local appearance purification (LAP), which improves generation quality and makes encoders such as MAE effective alignment targets by removing locally predictable content and aligning to the residual. LAP modifies only the alignment target and requires no changes to diffusion training or inference. Through controlled interventions within a fixed encoder, we show that effective alignment depends on what information the target emphasizes, not simply how much image information it contains. Our theory interprets standard alignment as maximizing a variational lower bound on the information that the denoiser's hidden state carries about the target, with purified alignment yielding a corresponding bound conditioned on clean latent patches. We further derive an exact constant-readout baseline that measures how well the target can be matched without using the hidden state. LAP improves generation across five pretrained representations, with the largest gains for MAE: on ImageNet-256 with SiT-B/2, it reduces FID from to . Its gains persist when combined with stronger alignment recipes and tokenizers. Across 13 encoders, the constant-readout score predicts the benefit of purification before diffusion training (Spearman ).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.