Heterogeneous-Negative Preference Optimization for Real-World Super-Resolution
Abstract
We present a new approach to explore Direct preference optimization (DPO) technique for real-world image supre-resolution (SR). First, we propose a preference-pair construction strategy in which the ground-truth high-quality image serves as the preferred sample, while predictions from a heterogeneous pool of existing SR models serve as negative samples. This strategy requires neither human annotations nor external reward models, while capturing diverse failure modes and providing more informative supervision for preference learning. Second, based on the constructed preference pairs, we derive a general DPO formulation applicable to arbitrary representation spaces. Further, we instantiate this formulation in the feature space of a pretrained DINOv2 model and introduce an adaptive layer-weighting mechanism to integrate representations across different semantic levels. Third, we decouple the shared temperature coefficient in DPO into two independent coefficients that separately control attraction toward preferred reconstructions and repulsion from negative predictions, thereby accounting for the asymmetric optimization difficulty of these two opposite directions. Extensive experiments on both synthetic and real-world datasets demonstrate the effectiveness and generality of the proposed framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.