acceptodds
Under review as a conference paper at ICLR 2027

Exploring the Design Factors of Pixel-Space Diffusion

Abstract

In this paper, we introduce EiT by exploring the key design factors of pixel-space diffusion from three perspectives: model architecture, representation, and frequency domain. Inspired by Denoising Autoencoders (DAE) which performs denoising through a bottleneck, we construct a U-shaped pixel diffusion transformer that performs explicit compression along the channel dimension. In terms of representation, we find that both pixel and latent diffusion transformers exhibit a three-stage representation dynamics across layers and timesteps: denoising/local feature expansion, basis lift-up, and posterior projection. Moreover, earlier denoising and a longer basis lift-up stage are associated with better generation, which EiT naturally achieves through its channel bottleneck and increased depth. From the frequency perspective, we establish the equivalence between pixel-space and frequency-space diffusion, and show that the latter enables flexible frequency-band loss reweighting, yielding performance gains at no extra cost. On ImageNet 256256 and 512512, EiT achieves significantly better generation quality than JiT (Just image Transformers). Our code will be made publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.