acceptodds
Under review as a conference paper at ICLR 2027

Unlocking the Potential of Contrastive Regularization for Diffusion Transformer Training

Abstract

Representation regularization has emerged as an effective way to improve diffusion transformer training. Yet one of the most canonical regularizers–InfoNCE-style contrastive learning–remains largely underexplored in this setting. We naturally ask a question: can standard contrastive learning serve as an effective regularizer for diffusion transformers? Surprisingly, we find that naive InfoNCE regularization brings only marginal gains. Our analysis reveals that the objective is easily dominated by a shortcut: near-orthogonal noise components in diffusion inputs provide a bypass for enlarging negative-pair distances while weakening positive-pair alignment, especially at high noise levels. To unlock the benefit of contrastive regularization, we propose lean epresentation-chored Contrastive () regularization. Specifically, anchors contrastive dispersion with clean-sample representation distances, transferring clean relational semantic structure to noisy diffusion representations to suppress the noise shortcut. Extensive experiments on ImageNet show that  consistently improves diffusion transformer training across model scales and training budgets, outperforming internal feature-alignment and dispersive regularization baselines while matching external DINO-guided REPA. These results re-establish InfoNCE-based contrastive learning as an effective regularizer for diffusion training, once its noise-driven over-dispersion is properly constrained.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.