Divide and Lighten: Lightweight Pixel Diffusion Transformers for One-Step Image Super-Resolution
Abstract
One-step diffusion substantially reduces the denoising cost of super-resolution, yet most existing methods still rely on latent diffusion architectures and thus remain burdened by expensive VAE encoding and decoding. We therefore propose LiPDiTSR, a Lightweight Pixel Diffusion Transformer that eliminates the VAE and performs one-step super-resolution directly in pixel space. LiPDiT-SR first builds a full VAE-free one-step PixelDiT model for super-resolution, where the patch-level and pixel-level pathways primarily handle semantic reasoning and texture reconstruction, respectively. We then derive a compact model by pruning the computationally dominant patch-level pathway while preserving the pixel-level pathway. Modulation Deviation Block Selection (MDBS) selects the patch blocks whose removal least disrupts the modulation signals delivered to the pixel-level pathway, while a Semantic Modulation Alignment (SMA) loss aligns these signals with those of the preceding reference model at each pruning stage. Experiments on multiple synthetic and real-world benchmarks demonstrate that LiPDiT-SR achieves visual quality comparable to state-of-the-art one-step diffusion SR methods while delivering a 1.6× inference speedup for ×4 super-resolution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.