PNR-SR: Policy-Native Reward-Policy Co-Training for One-Step Real-World Image Super-Resolution
Abstract
One-step generative models have substantially improved the inference efficiency of real-world image super-resolution (Real-ISR), while reward-based post-training provides a promising paradigm for improving restoration quality beyond supervised learning. However, existing approaches typically rely on an offline-trained reward model that remains fixed during policy optimization. As the SR policy evolves, its output distribution can progressively diverge from the reward model's learned preference distribution, weakening the reward model's ability to reliably distinguish restoration quality on newly generated outputs. In this work, we formulate reward-based post-training for Real-ISR as a closed-loop optimization problem and propose PNR-SR, a Policy-Native Reward-Policy Co-Training framework for one-step Real-ISR. To this end, we introduce HR-exposure ordering to derive policy-native preference supervision from the current SR policy, enabling online reward model optimization without human annotations or fixed external metrics. The reward model and SR policy are subsequently optimized in an online closed loop, where the reward model is continuously updated with policy-generated preferences to track the evolving SR policy distribution and provide dynamically updated reward guidance. To stabilize online reward model optimization, we further introduce dual-anchor calibration and a historical reward baseline to mitigate reward score drift. Extensive experiments demonstrate that our framework effectively improves one-step Real-ISR performance and outperforms existing state-of-the-art methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.