acceptodds
Under review as a conference paper at ICLR 2027

SNAP: Error Bars on Benchmark Averages Can Miss Run Covariance

Abstract

Error bars on a language model's benchmark average usually sum per-benchmark variances and ignore the covariance between benchmarks across training runs. On 375 DataDecide runs, the standard deviation of the ten-benchmark average in per-byte margin is 1.244 [1.143, 1.338] times its value under independence, while accuracy gives 1.078 with an interval that includes one. A margin comparison that assumes independence at a nominal 0.05 therefore has a true size near 0.11. That covariance sits within scoring formats, where BoolQ, alone in its format, carries 68% of the margin covariance trace. SNAP, a cross-half estimator, also gives per-benchmark run variances net of item noise, which the replicate standard deviation can't separate when item noise is large. A correction for the covariance changes few of DataDecide's recipe orderings. A pre-specified test on four held-out tasks fails at 1.217, with a lower limit below one. A second pre-specified rule also fails, since nine seeds at five PolyPythias sizes give margin inflation of 1.302 with an interval [0.921, 1.595] that includes one. A disjoint item bank leaves the source of that inflation undecided. All other analyses in the paper are exploratory. We recommend that authors report the replicate standard deviation of the battery average with a benchmark-removal check, and that standard deviation needs no item split and gives an inflation of 1.237 here.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.