Relative-Margin Reward for Candidate-set Ranking in Text-to-Image Generation
Abstract
Text-to-image models generate multiple candidates per prompt, and reward models are used to select among or rank them. Although benchmarks with human preferences over multiple images exist, reward model evaluation still centers largely on pairwise accuracy. Whether this metric adequately reflects a model’s ability to recover human rankings over entire candidate sets remains underexplored. We introduce Rank-Bench, a benchmark that recasts reward-model evaluation as tie-aware ranking over candidate sets, comprising 902 rankings over 6,236 images across 11 diagnostic dimensions. Rank-Bench reveals that comparable pairwise accuracy does not imply comparable set-level ranking agreement. We propose MarginRM, a relative-margin reward model trained to predict criterion-level signed ordinal margins and explanations under prompt-specific rubrics. At inference, it scores each candidate relative to a shared anchor under a common rubric, yielding a ranking over the full candidate set. Under all-pairs evaluation on Rank-Bench, MarginRM achieves 61.07% stable accuracy and a Kendall's of 0.519, compared with 54.47% and 0.487 for the strongest baseline. Single-anchor inference yields a of 0.485 using only 27.93% of the all-pairs comparison budget. As the reward for GRPO post-training of FLUX.1-dev, it improves GenEval from 63.72 to 72.61 and T2I-CompBench from 49.22 to 54.38, while producing rankings through shared-anchor comparisons without exhaustive pairwise evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.