Balancing Structure and Stability in Leaderboards with Quantal Response Equilibrium
Abstract
Public leaderboards shape how AI models are perceived and adopted. Yet defining a ranking from preferences is nontrivial, as different aggregation principles can favor different models. In practice, not every pair of models is compared, raising the additional challenge of predicting missing entries while retaining competitive structure. Elo’s scalar model restricts comparisons to additive logits, excluding potentially non-transitive structure. On frontier leaderboards, we show that adding a synthetic model that loses to everyone dethrones Elo’s leader, despite that leader still beating every rival and all original comparisons remaining unchanged. Maximal lotteries (MaxL), an alternative from probabilistic social choice, capture arbitrary tournament structure, but require a complete comparison matrix. Their outputs can change substantially across resamples. To reconcile structure and stability, we decouple leaderboard construction into comparison-matrix estimation and ranking. For estimation, we derive error bounds for low-rank completion that account for structural approximation, regularization, and finite-sample error. These bounds quantify the statistical cost of richer representations. For ranking, we analyze quantal response equilibrium (QRE), an entropy-regularized counterpart to MaxL that yields smooth, payoff-based scores. We derive temperature-dependent stability bounds and approximate guarantees for Condorcet consistency and independence of irrelevant alternatives. Empirically, QRE preserves the incumbent order in the insertion example and retains its leader in nearly all resamples.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.