acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Pixel-Space Diffusion Transformers: From Depth-Wise Connectivity to Layer-Wise Attention Patterns

Abstract

Recent studies have demonstrated the potential of Diffusion Transformers (DiTs) in pixel space. However, pixel-space models still suffer from slow convergence, high training costs, and the lack of native training at high resolutions. In previous work, fine-tuning a 1.0B text-to-image model at resolution in pixel space required industrial-scale cluster resources. In this work, we diagnose two potential factors underlying the inefficiency of pixel-space DiTs: depth-wise connectivity and layer-wise attention patterns, and propose the Adaptive Connection Image Transformer (AceiT). AceiT achieves reliable convergence with as little as 7 GFLOPs per sample, reducing the computational cost to one quarter of JiT-B/16 and demonstrating substantial gains in computational efficiency. Meanwhile, AceiT achieves faster convergence, alleviating the reliance on representation alignment (e.g., REPA) while remaining free from additional network components (e.g., pixel decoders), thereby offering better simplicity and scalability. More importantly, with only A100 GPUs, AceiT can be pretrained at resolution with up to 2.1B parameters, making native high-resolution training from scratch in pixel space feasible. We will release our code to facilitate further research and development in the community.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.