When Does a Benchmark Score License a Capability Claim? An Executable Three-State Identification Certificate
Abstract
A benchmark score is routinely read as evidence that a model has a capability, yet the same score can be produced by a non-capability side channel: acquisition metadata, option form, question templates. We ask when a score licenses its capability claim, and turn the question into an executable procedure. Given a claim, a frozen model score S, and a pre-declared nuisance class T , our certificate returns FAIL when a witness inflates the score without adding intended evidence, SURVIVE when no admissible witness moves it beyond a pre-registered equivalence margin, and INCONCLUSIVE when the data cannot separate the two at that margin, reported together with the inflation the data can resolve. A SURVIVE is earned by a two-one-sided-test bound calibrated against injected channels of known strength, so “no witness found” can never be mistaken for “benchmark is valid”. A companion identity delimits scoring-side repair: the maximum effective sample size of any channel-neutralizing reweighting equals a mass-weighted overlap efficiency, so localized channels are cheap to repair, while the corpus-wide near-deterministic channels our failures exhibit can only be disclosed. Applied to seven audit units, six of them the benchmarks’ own published headline claims, the verdict changes conclusions in both directions: two claims fail, two are positively preserved, and three are returned INCONCLUSIVE and then separated by a power analysis rather than left open. The certificate is a measurement instrument, not a benchmark-condemning machine.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.