Validating Rankings from Pairwise Comparisons
Abstract
Pairwise human preferences are widely used to evaluate and rank large language models on open-ended tasks, as in Arena chiang2024chatbotarena. This raises a natural question: how can we determine whether a reported ranking is good? More specifically, which measures of ranking quality can be reliably validated when the ground truth is unknown and only finite comparison data are available? We formalize this question through statistical justifiability of an error metric, asking whether there exists a validation algorithm that can reliably distinguish outputs with small error from those with substantially larger error using finite comparison data. We first show that ranking error, a natural metric for the central objective of ranking accuracy, is not statistically justifiable. This motivates seeking alternative, statistically justifiable error metrics. We develop metrics based on tests with observable outcomes, tailoring the tests to the information available for validation. For example, when the ranking method also returns estimated probabilities that one model is preferred over another, these enable pairwise multicalibration tests, which assess agreement with observed preferences within slices of the data such as tasks or domains. When only a reported ranking is available, we measure its compatibility with calibrated predictions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.