acceptodds
Under review as a conference paper at ICLR 2027

Ordered Representations for Pixel Diffusion: Denoising with Learned Velocity Correction

Abstract

Recent pixel diffusion models improve generation through clean-image prediction (-pred) or encoder–decoder architectures that reduce the noise-modeling burden on the backbone. Combining these designs, we find that -pred learns stronger representations and achieves lower FID than velocity prediction (-pred) with sufficient sampling steps, yet underperforms in few-step sampling. Examining its induced sampling velocity reveals that the -to- conversion amplifies prediction errors, which are most pronounced along the leading, high-energy directions of RGB images. Motivated by this structure, we propose **OiT**, an Ordered-representation image Transformer that uses nested channel masking to make dominant image information progressively accessible from the leading feature channels. A compact prefix of this representation guides a full-dimensional velocity correction, while the full representation remains responsible for clean-image prediction. This allows OiT to improve the representation quality and long-step generation of -pred while achieving the strong few-step performance of -pred. We further introduce bottlenecked Representation Alignment (bREPA) and a Hann Decoder as lightweight designs for representation alignment and patch decoding. On ImageNet , OiT-B and OiT-XL achieve FID-50K scores of 2.52 and 1.55, respectively, reaching state-of-the-art performance among pixel diffusion models at comparable compute.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.