acceptodds
Under review as a conference paper at ICLR 2027

Attract to Align, Repel to Differentiate: Contrastive Representation Alignment for Diffusion Models

Abstract

Representation alignment effectively accelerates diffusion transformer training by pulling hidden states toward frozen vision model features. However, solving independent point-to-point regression makes this alignment purely attractive. It lacks the repulsive force needed to differentiate spatially distinct patches, causing inefficient structural transfer and feature oversmoothing. We propose Contrastive Representation Alignment (CoRA) to mitigate representation collapse by introducing an intra-image contrastive objective. By explicitly contrasting student patches against distinct teacher patches, CoRA provides the missing repulsive gradient to maximize spatial micro-diversity. To prevent destructive gradient interference, we decouple the alignment and repulsion objectives into parallel pathways. This enables safe scaling of the negative pool size and ensures seamless compatibility with mainstream frameworks, including convolution-based approaches like iREPA. Extensive experiments demonstrate that CoRA consistently accelerates convergence and improves generative quality, adding negligible training overhead and zero inference cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.