acceptodds
Under review as a conference paper at ICLR 2027

Leaderboards rank finer than they resolve

Abstract

Machine learning adjudicates progress by ranking methods on shared benchmarks. But a benchmark is an instrument, with a resolution below which it returns noise. That resolution is almost never reported, so a reader cannot tell which leaderboard gaps are evidence and which are chance. Power analysis names it the minimum detectable difference: the smallest true accuracy gap a benchmark's own significance test detects at least eight times in ten. We ask three questions of it. First, how much of what the field publishes does its data support? Across 202 leaderboards in eight families, the median board ranks dozens of systems from first to last, yet its data can separate them into only two to five groups, depending on how strictly a group is drawn. Second, does a benchmark's resolution predict which of its published orderings hold up when measured again? No resolution measure has been checked against a second measurement of the same ordering. We rescore each ordering on items disjoint from those that assigned its label. Orderings the benchmark cannot resolve reverse in 21.5% of cases, against 0.27% of those it can, and the same test separates on four further boards, one in an unrelated field. A benchmark's resolution tells a reader, before any replication exists, which orderings to trust. Third, is resolution a property of the benchmark alone, so that its authors could compute it once and print it on the benchmark's card? It is not: it also depends on the rate at which the two compared systems disagree, which belongs to the comparison rather than to the benchmark, and a rate borrowed from a neighbouring board keeps its shape but loses its scale. Resolution has to be measured on the board itself, from the per-item outcomes its leaderboard already holds. Doing so costs no experiment, and leaderboards should report it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.