Pair2Point: Reliable as Pairwise, Efficient as Pointwise Rewards for High-Quality Text-to-Image Generation
Abstract
Recent multimodal reward models (RMs), especially those built on powerful vision-language models, have shown great promise in aligning vision models with human preferences. However, they face two main challenges. First, their training images often come from earlier generations of text-to-image models and no longer match the quality of current generators. Second, it is difficult to combine the reliable judgments of pairwise comparison with the efficiency of pointwise scoring. To address these challenges, we propose Pair2Point. We build three resources with recent generators for training and evaluation, including P2P-CoT-16K, P2P-Image-70K, and P2P-Bench. On top of these resources, a two-stage framework first trains a pairwise model with structured comparison labels from Gemini 3.1 Pro, so that it evaluates two images along relevant dimensions and explains its judgment; it then distills this model into a pointwise model through multi-opponent aggregation. Experiments show that Pair2Point achieves the highest accuracy among the evaluated dedicated reward models across multiple reward benchmarks. When used for down-stream reinforcement learning, it consistently improves bothtext–image alignment and visual quality. The resources and model weights will be released publicly after review.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.