acceptodds
Under review as a conference paper at ICLR 2027

(In)transitivity: Where Bradley-Terry Fails to Model Human Preferences

Abstract

Most preference-based alignment methods use Bradley-Terry (BT) preference models, which implicitly assume human preferences are consistent and transitive, i.e., that they can be collapsed onto a single scalar. In reality, however, heterogeneous human preferences can cause intransitivity through cyclic majority relations both at the user, as well as population level—a problem well-studied in social choice theory. While recent game-theoretic alignment methods propose cycle-aware objectives to handle this, whether such cycles matter empirically remains unclear. In this work, we answer this question by first auditing 9 widely used human-preference datasets and find that by construction, most lack the overlapping, repeated judgements required to analyze population-level intransitivity. We then construct INTRANSIT, a controlled dense preference dataset, through preference-completion methods that estimate the missing population-level preferences. We use this as a measurement tool for widely used preference models and show how performance degrades to near-chance agreement on intransitive samples. At the same time, the denser preference structure in the completed datasets contains useful signal in that preference models trained on this data perform as well as models trained on fully complete pairwise annotations. Finally, at inference time, using trained preference model outputs reduces regret and majority-rejected selections, and reinforcement-learning systems trained with preference models over this show improvements for certain preference objectives. Our results identify two challenges for preference alignment: collecting and measuring preferences to allow intransitive preferences to surface, as well as building objectives for the representation of these preferences without imposing a scalar ordering.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.