The Comparison Graph Determines What RLHF Can Learn: Spectral Identifiability, Error Bounds, and a Pre-Training Diagnostic for Preference Data
Abstract
Reward models for **RLHF** are trained on pairwise human comparisons, and the field treats that data as a quantity to be increased. We show the binding constraint is its structure, measurable before training. A preference dataset induces a *comparison graph* on prompt–response pairs, so reward learning under Bradley–Terry is estimation on that graph; lifting it through the reward model's feature map yields one matrix, computable in a single pass, whose spectrum settles which reward directions the data identifies, how fast error falls, and which responses can be scored reliably. We characterise the identifiable directions, bound finite-sample error pointwise by an effective resistance, and give a minimax lower bound with the same dependence on conditioning, so the limits bind every estimator. Two consequences are properties of the graph alone and hold at every scale and under every feature map. First, none of six widely used public datasets contains a single comparison between responses to *different* prompts, and each fragments into to components, so every reward component varying across prompts is unidentifiable at any dataset size, model size or embedding, including inside two trained reward models' own representations. This is a gauge that within-prompt **RLHF** may ignore and that pooled best-of-, reward-thresholded filtering, a shared value baseline and any reported mean reward may not: on the audited corpus, two rewards the data cannot tell apart, with identical log-likelihood in double precision, agree on 1.000 of within-prompt decisions while a pooled top- selection overlaps by only 0.385 and prompt ranking by mean reward moves at Spearman +0.728. Second, where the feature dimension no longer binds the rank the unlearnable subspace closes, but per-response conditioning spreads by to : the constraint moves rather than lifts. Linear, margin, **MLP** and transformer heads return bit-identical estimates along a direction the data cannot identify, and varying only topology at fixed budget changes error by , a gap that does not close over a range of budgets. We also report where our own selection objective fails, and argue the diagnostic rather than any selection rule is the contribution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.