acceptodds
Under review as a conference paper at ICLR 2027

Spectral-Aware Adaptation for Fidelity-Preserving Diffusion Super-Resolution

Abstract

Diffusion-based methods have recently advanced real-world image super-resolution (SR) by leveraging powerful generative priors, typically adapting frozen backbones via low-rank updates such as LoRA. Existing approaches assign LoRA rank uniformly across layers, implicitly assuming homogeneous structural importance. We show that diffusion UNets instead exhibit pronounced intrinsic rank heterogeneity: effective rank is lowest at the input and output extremities of the network and rises toward the bottleneck. Motivated by this observation, we propose SAFA (Spectral-Aware Fidelity Adaptation), a parameter-efficient adaptation framework that estimates the effective rank of every candidate module via singular value decomposition, keeps the most spectrally expressive ones, and distributes a fixed trainable-parameter budget among them in proportion to effective rank, without adding any inference-time cost. Compared to uniform allocation on the same backbone, SAFA concentrates capacity in structurally expressive layers and produces 2.27x stronger total update energy on RealSR while training 6.1x fewer LoRA parameters; against prior diffusion-based SR methods it improves fidelity metrics (PSNR, SSIM) and reference-based perceptual similarity (LPIPS, DISTS, FID) on DIV2K, RealSR, and DRealSR. Qualitatively, SAFA holds structure together in regions that are easy to hallucinate text, repetitive patterns and fine textures where prior diffusion-based SR methods introduce geometrically inconsistent detail. We find that this fidelity-oriented allocation trades off some no-reference perceptual scores relative to methods tuned for aesthetic realism, and we report this trade-off explicitly rather than claiming uniform improvement. We further verify that the heterogeneity is not checkpoint-specific: the same macro-pattern holds across SD1.5, SD2.1 and SDXL, where decoder layers average 1.8x the effective rank of encoder layers, although the precise spatial profile is architecture-dependent. Our results suggest that pre-adaptation spectral structure is a lightweight, gradient-free signal for guiding parameter-efficient adaptation, and that it transfers across diffusion UNet backbones.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.