acceptodds
Under review as a conference paper at ICLR 2027

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory

Abstract

Reinforcement Learning from Human Feedback (RLHF) can violate fundamental social-choice axioms, including Pareto optimality, majority consistency, pairwise-majority consistency, and Condorcet consistency. Prior work relates these failures to the aggregation rule implicitly implemented by reward learning: unrestricted Bradley–Terry training recovers the Borda count, which satisfies Pareto optimality but not the other axioms. Restricting the reward model does not generally resolve the problem; linear and any fixed-degree polynomial reward models can fail even Pareto optimality. This appears at odds with the empirical success of RLHF and motivates us to ask when its axiomatic guarantees can be recovered. We first show that, with unrestricted rewards, replacing comparison frequencies by majority labels changes the induced rule from Borda to Copeland, thereby recovering all four axioms. For linear rewards, we characterize the regularized estimator as a projection of its unrestricted counterpart and identify how feature geometry and regularization determine whether these guarantees survive. For polynomial rewards, we establish a sharp distinction in the order of quantifiers: although every fixed degree admits an adverse candidate set, every fixed finite candidate set has an interpolation degree beyond which unrestricted aggregation is recovered. Above this threshold, standard RLHF recovers Borda and hence Pareto optimality, whereas majority-based RLHF recovers Copeland and satisfies all four axioms. Together, these results show how aggregation design and reward expressivity jointly determine the axiomatic behavior of RLHF. Experiments on UltraFeedback and Chatbot Arena confirm this division of labor: the gap between the two label choices opens only as a reward model's capacity approaches the per-prompt unrestricted fit, and practical reward models remain far from that regime.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.