acceptodds
Under review as a conference paper at ICLR 2027

Benchmark Visibility from Annotations Alone: An Auditable Lower Bound on Blind Exploitability

Abstract

Whether a benchmark can see the capability it is cited for is a property of the corpus, and the field establishes it by running systems that ought to fail: late, expensive, and one corpus at a time. We show that much of the same information is computable from the annotations alone: no model, no judge, no frames, which matters because every corpus here keeps its media behind a signed agreement and its annotations public. The instrument is two counts and a curve. Announcement counts items whose gold answer is stated in the question; the blind bound is the best accuracy of an enumerated family of strategies that see the question's form and the answer distribution and nothing else. It is one-sided (it condemns and can never acquit, for a reason no instrument of this kind escapes), and announcement is a detector of the same failure rather than independent evidence, so the two counts must not be added, yet they flag nearly disjoint items. The curve is what the field is missing: the same estimator returns a positive value when the template carries no information at all, a null-induced baseline fixed by the group-size profile alone, and it collapses as the template isolates items; we construct a corpus on which it rises first instead, and which shape a given corpus takes is measured rather than deduced. Either way the same corpus reads very differently under two choices papers rarely state. Three common estimators mistake a corpus's answers for its structure by one mechanism: in-sample scoring and the ordinary nonparametric bootstrap leak the held-out answer back into its own evidence, and leave-one-out inside a group of two fits the single other item. The bootstrap is the dangerous one, because it inflates the bound substantially while looking like rigour, and deduplicating its own draws returns it almost exactly to the point estimate; a fit-and-test split is the repair. We validate the instrument against a documented defect on two corpus versions identical outside the repair, against a third-party repair its own authors performed, and behaviourally on an open-answer corpus, where the separation survives an intervention that moves the surface feature the mask is built from while leaving the question's content alone. Across the corpora we survey, 3 of the 5 option-regime corpora carry a statistically supported blind-exploitability signal under a family-wise max-statistic permutation test that prices the rule family's own selection into the null.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.