GRASP: Grounded Restoration with Agent-Refined Semantic Prompts for Real-World Image Super-Resolution
Abstract
Real-world image super-resolution (Real-ISR) requires preserving structures supported by a low-quality (LQ) input while synthesizing realistic fine details not directly observable in it. However, existing generative Real-ISR methods often leave the complementary roles of visual and semantic conditions implicit, limiting control over structural fidelity and detail realism. To address this challenge, we propose Grounded Restoration with Agent-Refined Semantic Prompts (GRASP), a Real-ISR framework that explicitly assigns complementary roles to local visual evidence, global image context, and semantic prompts. Specifically, we introduce a Decoupled Prior Injection mechanism that separates local visual evidence from global image context. Within this mechanism, a diptych provides spatially aligned local evidence to anchor observable structures and guide local detail synthesis, while a lightweight Vision-Context Bridge (VCB) supplies global image context to stabilize the scene-level layout. We further propose an inference-time agent-driven prompting strategy that iteratively refines semantic prompts to guide realistic detail synthesis in ambiguous regions. Experiments show that GRASP achieves strong perceptual performance across multiple Real-ISR benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.