The Shortlist Illusion: Why Leaderboard Separation Fails Candidate Models
Abstract
Public LLM leaderboards rank hundreds of models across macro benchmark frames, and practitioners routinely rely on these aggregate scores to guide high-stakes model selection and enterprise deployment among a small shortlist of candidate models. We show that macro-level benchmark performance fundamentally fails to transfer to candidate shortlists, creating a pervasive shortlist illusion. Using cohort-conditional separation—the fraction of within-cohort model pairs distinguishable beyond an operational performance margin—we audit 9,103 model–date records for 4,108 models across 133 snapshots of the Open LLM Leaderboard v2. We isolate three distinct failure mechanisms that misinform deployment decisions: (1) Non-identification: separation is fundamentally undefined without naming the comparison cohort; on the same MUSR snapshot, a 35-model frame separates 69.6% of pairs, whereas a 5-model candidate shortlist separates only 20.0%. (2) Selection-induced compression: filtering top- models restricts the observed score range, collapsing discriminative power; on MATH-Hard, separation plummets from 91.1% across the full frame to 0% among the top 5 models. (3) Small-sample noise: while random sub-cohorts are unbiased in expectation, small candidate sets suffer severe dispersion, reducing top-model selection to near-chance agreement. We translate these findings into an actionable, zero-cost six-field reporting contract and recommend cohort conditionality as an indispensable auditing requirement to curb misleading model cards and prevent wasteful deployment choices.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.