What Is the Top of a Leaderboard Worth? Measuring and Correcting Selection Bias in Tabular and Time-Series Benchmarks
Abstract
A leaderboard's top row is a selected quantity: the winner minimizes average error over the very datasets on which its margin is then reported. We measure what this costs on four method dataset matrices (TabArena, GIFT-Eval, fev-bench, and a long-horizon grid of our own) with a subsampling audit whose ground truth comes from disjoint held-out datasets. At , the suite size typical of long-horizon forecasting tables, TabArena's top entry keeps only % of its reported advantage over a tuned gradient-boosting baseline out of sample on our primary scale (% to % across three aggregations, so is the defensible ceiling); the reported number one is the true number one % of the time; on GIFT-Eval the leader's margin over the runner-up reverses sign; and drawing whole source families rather than i.i.d. units costs more still. Textbook remedies fail: leave-one-dataset-out returns exactly the naive value in % of folds on all four matrices, and the classical envelope overpredicts measured optimism by -, mostly because entries differ in true effect rather than because they correlate (- against -). We therefore introduce SCB, a selection-corrected bootstrap that resamples datasets jointly across methods around a variance-matched plug-in, needs only the published matrix, and has a cluster-blocked variant for families. It removes bias where there is bias and does not cry wolf where there is none: at full suite size TabArena's margin gives up about a third, while two champions and a published ICLR 2025 table keep % or more. Code, data and corrected leaderboards are in the anonymized supplement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.