PhaDiSR: Phase-Aware One-Step Diffusion Transformer for Real-World Image Super-Resolution
Abstract
Real-world image super-resolution increasingly benefits from the strong generative priors of diffusion transformers (DiTs), yet one-step DiT super-resolution remains prone to periodic, grid-like artifacts. Using a translation-equivariance diagnostic, we trace these artifacts to token-phase dependence: prediction discrepancies in packed DiTs vary periodically with the packing stride, whereas a convolutional U-Net counterpart shows no such periodicity. Motivated by this diagnosis, we propose PhaDiSR, a phase-aware one-step diffusion transformer for real-world image super-resolution that addresses token-phase dependence through velocity correction and wavelet discrimination. At the prediction level, Polyphase Token Reassembly exposes the implicit token phases in the packed velocity prediction, models their local spatial relations, and reassembles a phase-aware correction before the endpoint update. To penalize the resulting artifacts, Token-Phase Wavelet Discriminator uses multi-scale, multi-shift Haar supervision to effectively capture diverse local token-phase groupings. To facilitate training, Self-Prior Projection uses the frozen initialization prior to construct projected training references. PhaDiSR eliminates periodic grid artifacts and achieves leading perceptual fidelity across standard benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.