Manufactured Curl in Preference Learning
Abstract
Scalar reward models underpin reinforcement learning from human feedback (RLHF) for large language models (LLMs). But do cycles in human preferences imply that scalar rewards are inadequate? We examine how preference aggregation, prediction targets, and judgment dependence shape the answer. Averaging choice probabilities can create nonadditive log-odds even when every annotator follows scalar utilities. Concentrating disagreement on one response can make this distortion substantially larger. Recovering mean utilities and predicting choices can favor different estimators; calibration and smoothing affect their comparison. Observation design matters independently: shared rankings and independent judgments can have identical Bradley–Terry probabilities but threefold different asymptotic variances in estimated curl. Reassembling human summarization judgments across 363 contexts into disjoint rather than shared worker groups increases mean absolute curl by 24.4% at matched counts. Simulations at the observed comparison counts quantify detection limits, while held-out prediction tests sensitivity to regularization. These results show why measured preference cycles require an audit of aggregation and observation design before they can justify replacing scalar reward models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.