Reject to Align: Selective Sample Rejection for REPA
Abstract
Representation alignment (REPA) has become a standard technique for accelerating Diffusion Transformer (DiT) training by regularising hidden states toward features from a pretrained encoder. However, REPA computes its alignment gradient as a plain arithmetic mean over the batch, assigning equal weight to every sample regardless of alignment quality. We use a variational analysis to relate the population REPA objective to a Barber–Agakov lower bound on the conditional mutual information between DiT and encoder features. This analysis interprets cosine similarity as a positive-pair compatibility score and motivates examining per-sample alignment gradients. The gradient analysis characterises how alignment updates vary with cosine similarity. Complementary empirical measurements link rejected samples to weaker alignment and a higher alignment-to-denoising gradient-norm ratio. These observations motivate using large per-sample loss deviations as a practical indicator of atypical alignment updates. Based on this analysis, we propose Reject-REPA (REPA), a selective sample rejection mechanism that filters alignment-loss outliers via a -based threshold grounded in the Vysochanskij–Petunin inequality, a dual detector combining Z-score + MAD for outlier identification, and cross-batch EMA statistics for adaptive thresholding. REPA introduces no additional model parameters and serves as a drop-in replacement for standard REPA training. Experiments on different generation tasks suggest that selective alignment can improve generation quality in the evaluated settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.