HiCoSR: High-Fidelity Single-Step Real-World Super-Resolution via Hierarchical Latent Diffusion
Abstract
High-fidelity real-world image super-resolution (RealSR) remains challenging for Latent Diffusion Models (LDMs) due to the fixed spatial compression in pre-trained variational autoencoders (VAEs), which causes fine structural information loss during encoding and often introduces artifacts in generated images. To address this limitation, we propose HiCoSR, a hierarchical and distribution-consistent framework designed to preserve spatial details in the latent space. Specifically, we design a HiCo-Encoder that introduces a parallel latent branch to maintain richer spatial and structural information beyond the capacity of the representation. Based on this dual-scale representation, we adopt a coarse-to-fine cascaded generation strategy to overcome the artifacts and semantic inaccuracies inherent in single-scale denoisers. The coarse stage first establishes a robust global layout, while the fine stage refines local structural details with a single denoising step at each stage. Extensive experiments on three RealSR benchmarks show that HiCoSR achieves strong reconstruction fidelity and comparable perceptual quality to other diffusion-based models, using only a pretrained SD-Turbo backbone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.