acceptodds
Under review as a conference paper at ICLR 2027

CORE: Learning Contextually Grounded Representations for Pixel-Space Diffusion

Abstract

Recent self-distillation methods improve latent-space diffusion by aligning noisy representations with cleaner targets, often derived from an exponential moving average (EMA) teacher. However, these gains do not consistently transfer to pixel-space generation, where models must learn spatial structure and fine-grained appearance directly from noisy pixels. We find that high feature similarity can persist even when spatial correspondence deteriorates sharply, revealing that representation alignment alone does not ensure utility for pixel prediction. To address this limitation, we introduce **CORE**, a spatial self-supervision framework for learning contextually grounded representations in pixel-space diffusion. *Correlated Dual Observations* provide paired noisy and cleaner views of the same image with correlated noise for training a shared generator. Built on these observations, *Contextual Representation Grounding* uses two coupled prediction tasks on a shared representation at masked locations. One predicts cleaner features from visible context, while the other predicts correction targets formed by selecting and scaling cleaner-induced prediction changes according to ground-truth pixel residuals. Together, these tasks encourage contextual recoverability while linking representation learning to pixel prediction errors. CORE requires neither an external representation encoder nor an independent EMA teacher and leaves inference unchanged. Experiments on ImageNet at and demonstrate that CORE achieves consistently lower FID than the evaluated self-supervision baselines. At 600 epochs, CORE-L/16 even achieves a lower FID than PixelREPA, which uses external representation supervision. Code will be publicly released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.