How Many Votes Make a Valid Leaderboard?
Abstract
Pairwise-vote leaderboards rank language models, but their displayed distinctions may exceed what the votes establish. We ask which leaderboard claims are statistically supported and which comparisons would resolve the remaining uncertainty. Our framework treats the best model, a top-k set, and a quality–cost frontier as separate claims. Under the Bradley–Terry model, each claim is tested against parameter regions supporting a different answer. We measure evidence by the smallest likelihood loss required to reach any such region. Tractable convex decompositions and conservative covers make these tests computationally feasible. We derive an anytime-valid stopping rule controlling false certification under adaptive sampling and stopping. To select new comparisons, we approximate the likelihood loss through the comparison graph’s information geometry. Our algorithm, CERTIFY, targets comparisons that strengthen evidence against the hardest remaining alternatives. An audit of 1.16 million Chatbot Arena votes across 129 models reveals a gap between stability and certification. Split-half rankings remain highly correlated, yet none of the 63 adjacent top-64 positions clears our proved threshold. Even 355,507 comparisons among the top 20 provide only 84% of the evidence required to certify the leader. Thus, a reproducible ordering need not establish its individual positions. On held-out recorded votes for a 16-model subset, CERTIFY identifies the leader in a median 12,480 comparisons. Across 15 runs at a nominal 5% error level, every returned answer is correct. Under the same stopping rule, competing sampling strategies require approximately 5.6–7.5x as many votes. Experiments also examine quality–cost and two-attribute frontiers, extending the approach beyond identifying a single winner. Together, these results connect the statistical support for leaderboard claims to the comparisons needed to establish them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.