acceptodds
Under review as a conference paper at ICLR 2027

SURE: Evolving Pairwise Reward Models with Rubrics from Uncertain Comparisons

Abstract

Pairwise reward models (RMs) are widely used as scalable proxies for human preferences in language model alignment, representing the relative quality of two responses through their reward difference. However, their judgments can become unreliable under distribution shift, particularly for comparisons near the model's decision boundary or requiring criteria and domain knowledge insufficiently covered during training. We propose \SURE, a sample-efficient framework for evolving pairwise RMs with rubrics learned from uncertain comparisons. SURE identifies informative comparisons from an unlabeled target set using two complementary signals calibrated against in-distribution anchors: reward-margin uncertainty and representation-space deviation. It queries a stronger evaluator only for the selected pairs and uses the resulting judgments to re-train the RM, while deriving hierarchical rubrics containing general principles and domain-specific criteria from the accompanying rationales. The rubrics are further refined through Group Relative Policy Optimization (GRPO)-inspired textual optimization, which converts group-relative agreement with the stronger evaluator into natural-language feedback rather than parameter gradients. At inference time, the evolved rubrics provide explicit criteria for preference judgment, complementing the adapted RM. Experiments show that, under matched selection ratios or external-query budgets, SURE identifies more informative uncertain cases, improves pairwise preference accuracy under distribution shift, and delivers reasonable gains on downstream tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.