Scale as a Modality: Multi-Scale Pixel-Space Diffusion Transformers
Abstract
Latent diffusion models rely on pretrained autoencoders that introduce lossy compression, decouple representation learning from generative modeling, and constrain the spatial granularity at which images are represented. Pixel-space diffusion removes this bottleneck and, importantly, enables flexible tokenization of the same image at multiple spatial scales. Existing pixel-space diffusion transformers, however, largely adopt two-level coarse-to-fine architectures with limited cross-scale interaction. We introduce MS-pDiT, a multi-scale pixel-space diffusion transformer that treats spatial scale as a modality. MS-pDiT tokenizes the same noisy image at multiple patch sizes, assigns each scale a dedicated transformer stream, and couples these streams through compacted MM-DiT-style joint attention. This design enables bidirectional information exchange across scales while keeping the attention cost independent of the finest patch size. It further establishes the number of scales as a new architectural scaling dimension. On ImageNet , increasing the number of scales consistently improves FID from 4.39 with two scales to 2.31, 2.26, and 2.19 with three, four, and five scales, respectively; under a comparable parameter budget, introducing an intermediate scale substantially outperforms additional depth. Our analysis further shows that bidirectional cross-scale interaction and late introduction of fine-grained streams are important for effective multi-scale modeling. MS-pDiT achieves FID scores of 1.43 and 1.44 on ImageNet at and , respectively, and naturally extends to text-to-image generation by incorporating text as an additional modality, attaining 0.86 on GenEval and 83.9 on DPG-Bench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.