Not Even Ordered: What a Published Leaderboard Entitles a Reader to Quote
Abstract
A leaderboard reports one total order over its methods, computed under evaluation choices its paper does not state. Those choices are a rule deciding when a prediction counts as the reference, crossed with an output budget, a post-processing step or a normalisation. We fix the items and the released predictions, re-score them in every cell of a declared grid, and ask what part of the published order survives. We do this for 24 leaderboards in four domains, retraining nothing and reimplementing no method. Together those tables print 298 places. Read permissively, a pair stands unless some cell contradicts it, and 170 ranks stand. Read strictly, a pair stands only where every cell separates it, and 99 survive. The distance between the two readings is mostly not the choices: 159 ranks go to resolving power a benchmark never had, before any convention is varied, and 40 to the grid. The surviving share runs from none to nine tenths, and does not replicate between two editions of one shared task. Survival is a property of a table and its grid, not of a benchmark. Of the orderings a table did establish, 6 are reversed by a declared choice and survive their own board's correction, and 5 of those survive a single correction over all 24 grids. None of the six is the matching rule's, the choice this literature discusses most. Deciding which orderings are at risk needs no pairwise statistic. We release the protocol, the audited split, the frozen predictions, the checker and the analysis code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.