acceptodds
Under review as a conference paper at ICLR 2027

Rubric-Residual Preference Geometry for RLHF Signal Selection

Abstract

Preference collection for RLHF can conflate stable disagreement with choices induced by the instructions given to raters. We introduce Rubric-Residual Graph Reward (RRG-Reward), a preference model that separates these two roles before using local residual agreement as a graph regularizer. The method cross-fits an arm-adjusted annotation model, builds edges only from pre-specified repeated co-labeling, and regularizes a latent preference mixture with the resulting graph. Its policy interface uses a reversal-antisymmetric, variance-normalized margin, so the same uncertainty term cannot reverse a pairwise preference when response order changes. We give matched equations for graph-free mixture and global-factor comparators, a target construction for independent binary labels, a single clustered inference estimand, and a direct human-evaluation aggregation rule. Together, these definitions make rubric-residual preference geometry a precise and testable signal-selection primitive for disagreement-aware RLHF.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.