acceptodds
Under review as a conference paper at ICLR 2027

TAPD-SR: Bridging the Detail-Conditioning Gap for Diffusion-Based Faithful Real-World Image Super-Resolution

Abstract

Real-world image super-resolution (Real-ISR) using pretrained text-to-image (T2I) diffusion models achieves strong perceptual quality but often suffers from unfaithful textures. We trace this failure to a detail-conditioning gap: loss of low-resolution (LR) features in latent space, under-specified by coarse textual prompts, and insufficiently enforced along the denoising trajectory. To address this issue, we present TAPD-SR, a unified framework that strengthens LR-derived conditioning to improve generated texture fidelity. At the representation level, a Time-Aware Detail Feature Extractor (TDFE) captures high-frequency information directly from the LR input and adaptively injects it into the diffusion process according to the denoising timestep. At the semantic level, a Reasoning-driven Fine-grained Prompt Generator (RFPG) produces fine-grained global and object-level prompts from the LR input using a fine-tuned multimodal large language model (MLLM). At the trajectory level, a Hierarchical Frequency Refinement Module (HFRM) extracts multi-scale frequency information from the LR input to predict a residual correction signal to the diffusion model's velocity output, providing additional structural and detail guidance for latent updates. Extensive experiments demonstrate that our TAPD-SR improves reconstruction fidelity while maintaining high perceptual quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.