Leaderboards Display Ranks, Benchmarks Certify Tiers: Simultaneous Inference and a Capacity Law for LLM Evaluation
Abstract
Public LLM leaderboards display a total order over models, but the finite item sample behind each leaderboard supports only a partial order. We make the certified objects explicit: simultaneous rank sets, the certified pairwise partial order, and certified tiers: ordered groups whose every cross-tier comparison holds jointly with familywise confidence . Auditing five public leaderboards across three data modalities shows that the data support a great deal (57–77% of pairwise orderings certified) while the displayed total order is almost entirely unsupported: AlpacaEval 2.0 shows 65 ranks but certifies 2 tiers; LiveBench 94 and 1; a 57k-battle Chatbot Arena release 63 and 2; SWE-bench Verified 134 and 1; the Open LLM Leaderboard's 4,576 ranks reduce to at most 2 tiers per benchmark (5 on its top-100 slice); and top-1 plausible sets range from a unique champion to 24 models (87 under task-clustered inference). A tier capacity law, an approximate scaling relation in the sense of empirical scaling laws, predicts how certifiable tiers grow with items as and collapse with model density (slope constants correct to within ; the crowding decline reproduced within 35–46% at the smallest subsets), and inverts into item budgets for target resolution. On two refreshing benchmarks, certified tiers did not invert on fresh items in the one refresh where the check is non-vacuous, and rank-flip rates match the closed-form prediction over 130 pooled adjacent pairs. We release tiercert, a CPU-only toolkit that turns any score matrix or battle log into rank sets, certified tiers, and item-budget prescriptions; code, data, and result files are provided as supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.