BandDiT: Band-Aligned Diffusion Transformers for Pixel-Space Image Generation
Abstract
Pixel-space diffusion models denoise pixel values directly, without the pretrained autoencoder that latent diffusion models use to compress images. At each denoising step, the pixel-space transformer must extract meaningful features from raw pixels, reason over these representations, and map its predictions back into image space. In contrast, latent diffusion models operate in a learned representation space that abstracts away some low-level detail, leaving the pretrained decoder to reconstruct fine details and textures. This places the full burden of spectral reconstruction on the transformer, raising a central question: how are different image frequencies represented across network depth, and how can this organization guide learning? We find that pixel-space diffusion transformers build the spectrum progressively across their blocks: low-frequency bands become decodable from shallow blocks, whereas high-frequency bands emerge near the output. This hierarchy lies in internal representations within a denoising step, distinct from the known coarse-to-fine evolution of samples along the sampling trajectory. We introduce \our, a band-aligned diffusion transformer that turns this hierarchy into explicit supervision. Wavelet bands of the clean image, which need no pretrained teacher, serve as targets for intermediate blocks at matching depths, from low frequencies in shallow blocks to high frequencies in deep ones. On ImageNet , \our-B/16 matches the baseline's FID and 5-crop FID with fewer epochs and reduces the high-frequency Fr\'echet Wavelet Distance by , indicating improved global coherence, local structure, and texture fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.