What Does Benchmark Capability Coverage Establish? An Audit of 16,217 Benchmarks
Abstract
New benchmarks are justified by gaps: no existing benchmark, it is claimed, tests some combination of capabilities. What does a coverage map of the benchmark ecosystem establish? We find that its conclusions depend on choices that maps leave implicit: how finely capabilities are distinguished, how many benchmarks must test a combination, what a benchmark's labels assert together, and which reference model defines “unusual”. We make these choices explicit in a coverage-audit instrument whose logic is checked in Lean 4: given benchmark assignments, a vocabulary resolution and a reference model, it reports supported combinations, relative deficits and their witnessing benchmarks. We instantiate it as an atlas of 16,217 benchmarks labelled with 754 capabilities and measure the labels against two human raters. Coverage is concentrated beyond what benchmark breadth and capability popularity explain: 564 benchmarks witness every covered coarse capability pair, against 812 under a margin-preserving null, and Papers-with-Code's own task tags show the same pattern. Which associations and modality restrictions look unusual changes with the reference model. The atlas yields an exploratory queue of deficit candidates with their witnesses. For one under-supported capability triple it found, we show that an executable benchmark can be built and diagnosed with demand-removal controls. Apparent gaps are audit targets, not certified absences. We release the vocabulary, labels, Lean development and pipeline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.