PairReFL: Pairwise Reward Feedback Learning for Diffusion Models
Abstract
Aligning text-to-image models to human preferences, guided by a reward model, has become common practice. Compared with policy-gradient methods relying soly on scalar reward, reward feedback learning methods backpropagate feedback through the reward model, extracting richer training signals from each sample. Despite its efficiency, this learning paradigm is susceptible to reward hacking: maximizing a reward model's per-image score can amplify undesirable visual patterns such as oversaturated colors, without improving perceived image quality. In this paper, we propose Pairwise Reward Feedback Learning (PairReFL), which replaces scalar-score maximization with a differentiable comparison against another image sampled from the current model, so the comparison reference changes as the model learns. Across four human-preference metrics, PairReFL outperforms pointwise reward feedback learning by up to 1.1% and the strongest policy-gradient baseline by up to 3.7%. Our differentiable pairwise reward model reduces the error rate in visual-quality comparisons by 62% relative to its pointwise counterpart. Further analyses trace these gains to the pairwise interface, which provides more stable feedback on closely matched samples and achieves greater preference gains for a given degree of model change.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.