Answer-Key Drift: Which Benchmark Conclusions Survive a Revised Key
Abstract
When a benchmark is revised, maintainers must decide which published model comparisons still hold. We study this on fixed model outputs, separating changes to answers, scored items, scoring rules and reference actions across twelve archived revision series, a text-to-SQL release and a 2026 tool-use benchmark revision. Headline results can stay put while specific conclusions move. On fixed arithmetic responses from 60 models, one answer-and-item update reverses 10, 9, 10 and 8 of 1,770 comparisons under four scoring rules, and 1,748 comparisons reverse under none of them; yet the four rules yield three distinct reversal sets, with no pair common to all four. On WorkBench, with scorer, tools and state fixed, updated reference actions for 9 of 679 tasks leave the champion and the top-five set unchanged but rewrite 179 of the 216 affected success judgments and reverse 4 of 276 comparisons. The revision's footprint also says what can be reused: a budget certificate that uses only the number of changed tasks certifies 252 of the 276 comparisons from the old scores alone, so all four reversals lie among the other 24. Replacing text-to-SQL references reverses at least six of 120 comparisons under set-equality scoring for every completion of failed evaluations. Across the archived series we express displacement in item-sampling units, locate where edits land, and compare budget rules with same-information baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.