acceptodds
Under review as a conference paper at ICLR 2027

Decouple then Revive: Towards Perceptual Super-Resolution via Reinforcement Learning with Anchored Generative Priors

Abstract

Diffusion models advance RealSR through progressive distribution matching, yet existing training paradigms remain challenged: supervised finetuning on paired data tends to produce over-smoothed textures or misaligned semantics, while RL-based approaches introduce structural and textural inconsistency. In this paper, we propose NA-DePO (Null-space Anchored Dense Preference Optimization), which decomposes the SR signal into observation-constrained and unconstrained components via null-space decomposition, and restricts RL exploration exclusively to the unconstrained generative subspace to leverage generative priors without compromising structural consistency. We decouple the optimization process: we maintain the LR fidelity via a range-space constraint, and conduct negative-aware contrastive policy learning within the null-space. Furthermore, to provide informative and stable guidance, we design a dense reward mechanism incorporating VLM-driven local reward maps and timestep-aware scheduling, effectively reducing multi-reward conflicts. Extensive experiments demonstrate that NA-DePO significantly improves the balance between realistic detail generation and faithful structural restoration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.