acceptodds
Under review as a conference paper at ICLR 2027

Purify Before You Align: Improving Representation Alignment for Diffusion Models

Abstract

Representation alignment improves and accelerates diffusion training by encouraging intermediate features to match a pretrained visual encoder. Yet choosing the right encoder remains largely empirical: DINOv2 provides substantial gains, whereas MAE, despite retaining more reconstructive information, offers little benefit. We introduce local appearance purification (LAP), which improves generation quality and makes encoders such as MAE effective alignment targets by removing locally predictable content and aligning to the residual. LAP modifies only the alignment target and requires no changes to diffusion training or inference. Through controlled interventions within a fixed encoder, we show that effective alignment depends on what information the target emphasizes, not simply how much image information it contains. Our theory interprets standard alignment as maximizing a variational lower bound on the information that the denoiser's hidden state carries about the target, with purified alignment yielding a corresponding bound conditioned on clean latent patches. We further derive an exact constant-readout baseline that measures how well the target can be matched without using the hidden state. LAP improves generation across five pretrained representations, with the largest gains for MAE: on ImageNet-256 with SiT-B/2, it reduces FID from to . Its gains persist when combined with stronger alignment recipes and tokenizers. Across 13 encoders, the constant-readout score predicts the benefit of purification before diffusion training (Spearman ).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.