Exploring the Design Factors of Pixel-Space Diffusion
Abstract
In this paper, we introduce EiT by exploring the key design factors of pixel-space diffusion from three perspectives: model architecture, representation, and frequency domain. Inspired by Denoising Autoencoders (DAE) which performs denoising through a bottleneck, we construct a U-shaped pixel diffusion transformer that performs explicit compression along the channel dimension. In terms of representation, we find that both pixel and latent diffusion transformers exhibit a three-stage representation dynamics across layers and timesteps: denoising/local feature expansion, basis lift-up, and posterior projection. Moreover, earlier denoising and a longer basis lift-up stage are associated with better generation, which EiT naturally achieves through its channel bottleneck and increased depth. From the frequency perspective, we establish the equivalence between pixel-space and frequency-space diffusion, and show that the latter enables flexible frequency-band loss reweighting, yielding performance gains at no extra cost. On ImageNet 256256 and 512512, EiT achieves significantly better generation quality than JiT (Just image Transformers). Our code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.