Balance of Benchmarks: Evidence Retention and Multiplicity-Aware Model Comparison
Abstract
Benchmark aggregation couples two quantities that should be set separately: how much evidence a comparison uses and how much influence each benchmark receives. Curated suites control influence by selecting benchmarks, potentially excluding informative observations, while equal weighting of a broader pool lets repeated or closely related benchmarks gain influence through sheer numbers. We introduce Balance of Benchmarks (BoB), which retains eligible observations, fits benchmark-specific response curves to place different difficulty ranges on a shared ability scale, and down-weights benchmarks in crowded semantic neighborhoods using description similarity. We evaluate evidence retention and density weighting through separate interventions on WildScores, a sparse collection of 148 developer-reported benchmarks, with 14 Artificial Analysis benchmarks as a controlled reference. Holding fitting models fixed and withholding source-lineage families, expanding one historical ten-benchmark suite to the full eligible pool raises BoB's held-out Spearman correlation by 0.035 over 50 targets (Holm p=0.013). Density weighting reduces sensitivity to complete re-listings: after four copies are added, mean absolute rank displacement falls from 1.95 under equal-weight equating to 0.63. Reporting masks reveal a further distinction: conserving a group's column weight leaves each model's observed weight dependent on its reporting coverage. Together, these results clarify how evidence retention, benchmark multiplicity, and reporting coverage jointly shape model comparison.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.