acceptodds
Under review as a conference paper at ICLR 2027

ViewCert: Verifying Model Selection under Evaluator Ambiguity

Abstract

Model selection is usually reported under one evaluator, even when the same public data admit several defensible choices. ViewCert verifies whether a model-selection claim holds for every program in an explicitly declared, mechanically complete typed language while reusing fixed outputs. Unstructured verification is coNP-complete. Factored margins reduce exactly to classical min-sum elimination, with cost exponential only in interaction width . ViewCert returns a replayable worst-case program, handles the full aggregation simplex without a grid, and jointly covers the program, competitor, weights, and budget selected after search. As a post-outcome algorithmic closure on Open Images, an audited width-two compiler verifies Grounding DINO Base against five competitors over evaluator programs without enumerating their leaves; its worst-case margin is AP. Prospectively on WMT25, five official metrics certify Shy over Gemini-2.5-Pro, while two registered LLM judges certify the reverse among 15,120 covered comparisons. Across six regimes, ViewCert distinguishes stable selection, certified ranking dependence, and principled non-identification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.