RELEVANCE-GUIDED TOKEN SUPERVISION ALLOCATION FOR DIFFUSION REPRESENTATION ALIGNMENT
Abstract
Recent work has shown that aligning intermediate diffusion features with pretrained visual representations can improve diffusion model training. However, existing methods typically use teacher representations to define alignment targets while allocating supervision uniformly, regardless of image content. To address this limitation, We propose Relevance-Guided Token Supervision Allocation (RGSA), which uses teacher-derived relevance cues to replace uniform aggregation of token-wise alignment losses with image-conditioned weighting.Specifically, RGSA derives proxy relevance cues from frozen teacher representations using patch-feature norms to characterize response strength and patch–CLS cosine similarity to capture local–global semantic consistency. These cues are fused into token relevance scores and normalized within each image to assign relative weights to the corresponding token-wise alignment losses. Experiments on ImageNet-1K show consistent improvements over iREPA across SiT-B/2, SiT-L/2, and SiT-XL/2, with similar gains across diverse visual teachers. Controlled ablations further show that RGSA outperforms random and spatially shuffled weighting in both FID and Recall. Together, these results highlight relevance-guided supervision allocation as an effective design choice for diffusion representation alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.