acceptodds
Under review as a conference paper at ICLR 2027

Comparing Models from Incomplete Evaluation Evidence

Abstract

Leaderboard tables assert an ordering over all models, but the evidence is incomplete: models are evaluated on different benchmark subsets. We ask which of those comparisons the observed evidence establishes. Taking the platform’s declared model universe, benchmark universe and aggregation operator as given, we add one counterfactual estimand: that operator on the matrix completed over that universe, unobserved cells held to admissible ranges. Auditing the reported aggregate would be tautological—it is that operator’s own output. Each comparison is labelled established, possible, refuted or undetermined. For HELM’s mean win rate, pairwise claims admit exact closed-form bounds under the pinned rule reconstruction, to which established leaderhood reduces; possible leaderhood is genuinely joint. On HELM’s classic: core_scenarios Accuracy universe (81.7% observed), 4 of the 45 reported top-ten orderings are established and no model leads in every completion; two further declared universes establish 0 and 3. Complete cases decide all 45 because a measured block leaves nothing to decide, while the two deletion remedies cost the same evidence but name different leaders. Masking a real block, the pattern of missingness, not its amount, changes what is established. The audit is also a decision problem, and one step of it is solved exactly: the measurement that settles the most reported claims whatever it returns is computable from the published matrix alone, is provably blind to the value it buys, and certifies in advance how many claims a budget will settle. :chatgpt-content-reference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.