acceptodds
Under review as a conference paper at ICLR 2027

RELEVANCE-GUIDED TOKEN SUPERVISION ALLOCATION FOR DIFFUSION REPRESENTATION ALIGNMENT

Abstract

Recent work has shown that aligning intermediate diffusion features with pretrained visual representations can improve diffusion model training. However, existing methods typically use teacher representations to define alignment targets while allocating supervision uniformly, regardless of image content. To address this limitation, We propose Relevance-Guided Token Supervision Allocation (RGSA), which uses teacher-derived relevance cues to replace uniform aggregation of token-wise alignment losses with image-conditioned weighting.Specifically, RGSA derives proxy relevance cues from frozen teacher representations using patch-feature norms to characterize response strength and patch–CLS cosine similarity to capture local–global semantic consistency. These cues are fused into token relevance scores and normalized within each image to assign relative weights to the corresponding token-wise alignment losses. Experiments on ImageNet-1K show consistent improvements over iREPA across SiT-B/2, SiT-L/2, and SiT-XL/2, with similar gains across diverse visual teachers. Controlled ablations further show that RGSA outperforms random and spatially shuffled weighting in both FID and Recall. Together, these results highlight relevance-guided supervision allocation as an effective design choice for diffusion representation alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.