Rethinking Fidelity–Perception Trade-off in Latent-Free Super-Resolution via Pixel Flow
Abstract
Real-world image super-resolution (SR) aims to reconstruct high-resolution images from low-resolution observations. Recent models using diffusion models and flow matching have shown strong perceptual restoration ability, yet we observe two key limitations that hinder their robustness and fidelity: 1) their performance is sensitive to different noise and sampling stochasticity; 2) per-step noise prediction may fail under composite degradations and accumulate residual errors, causing identity, texture, and structural drift. To address these issues, we rethink the fidelity–perception trade-off from a latent-free perspective and propose PixelSR, a pixel-flow SR framework that directly models the restoration trajectory in image space. Instead of predicting per-step noise in a compressed latent space, PixelSR predicts clean high-resolution images and derives the corresponding pixel-space flow for iterative refinement, thereby anchoring the generation process to the high-quality image manifold while preserving low-level structures from the input. This clean-data-centric formulation reduces error accumulation and provides a more stable path toward perceptually realistic yet faithful reconstruction. Interestingly, our experiments reveal that varying the number of sampling steps provides a simple and effective mechanism to navigate the fidelity–perception trade-off. Extensive experiments on synthetic and real-world SR benchmarks demonstrate that PixelSR achieves a better fidelity–perception balance than strong diffusion-based baselines, with improved identity preservation, structural consistency, and perceptual quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.