Sparse Agent Leaderboards: Identity, Support, and Identifiability
Abstract
Sparse agent leaderboards compare model-harness configurations under uneven coverage. A model ranking still needs one shared evaluation target. We map task identity, common-target support, and design identification onto three actions: certify a supported ordering, label a model-assisted extrapolation, or defer the comparison. Under task additivity, an incidence graph characterizes identifiable configuration contrasts, while free model-harness interactions leave model main effects unidentified. Sharp probability-scale bounds quantify what missing cells alone permit. In 22,818 archived Terminal-Bench trials spanning 53 configurations, candidate task-label joins connect a disconnected graph without adding any of the 4,702 observed cells, and a 55-record provenance sample contains four reported signatures within one name family. Under every label policy the best-covered model still lacks more than 60% of a uniform harness-task target, so no model pair can be certified against every bounded completion. On a separate public panel of 150 published aggregates, DeepSeek V3.2 beats GPT-5.2 under the source four-domain weights and loses under equal benchmark weights. At 90% observed cells the certificates cover 57.75-61.85% of pairwise orderings, while available-score ranking still publishes all ten pairs. Point estimates can match that coverage with almost no errors on this fixed table. The certificate is a worst-case column and a next-cell rule, not a rival accuracy score. Revealing ten additional deferred-weight cells on random 90% masks raises four-domain certificate coverage from 61.65% to 97.75%, versus 84.25% for ten random extra cells. A table-agnostic CSV tool returns the same three actions from observed cells.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.