You Don't Need to Run Every Eval
Abstract
A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare design choices, and select the checkpoint for the release. But do we need to run every eval? We compile a public score matrix of 84 frontier models on 133 benchmarks (2,604 cells, 23.3% filled) and find it is approximately rank-2: two factors capture most of how a model's scores vary across benchmarks. We confirm this in two ways: (i) scores hidden from the matrix are best recovered using two factors, and (ii) two factors already explain over 90% of the variation among models on the benchmarks they share. Building on this, we design BenchPress: a logit-space rank-2 matrix completion method with a median absolute error of 4.6 points on held-out scores. Using BenchPress, we find a subset of five benchmarks GPQA-D, HLE, Codeforces, MMLU-Pro, ARC-AGI-1 that can recover the rest of a model's public scorecard with a median absolute error of 4.74 points. For a tighter evaluation budget, a cheaper set GPQA-D, MMLU-Pro, Aider Polyglot, MATH-500, AIME 2026 can predict a model's remaining evals with a median absolute error of 5.32 points. Finally, we identify what affects prediction reliability and combine these factors with predictor disagreement to quantify when predictions can be trusted. Code and data are available at https://anonymous.4open.science/r/benchpress-iclr2027.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.