acceptodds
Under review as a conference paper at ICLR 2027

The Price of Choice in Selective Prediction

Abstract

A vendor may claim that some rule from a predeclared menu—a score threshold or a union of score intervals—lets its selective classifier keep a fraction of inputs while erring on at most a fraction of them. Refuting such an existential claim means ruling out every permitted rule at once. The intersection–union test (IUT) does this without a multiplicity correction, but its sample cost as the menu grows, and how close that cost is to the best valid audit, are not captured by the level argument. We characterize both. For a fixed interior risk target , menus of intervals and a risk gap , in the small-gap, large-menu regime of our theorem the minimax sample size at each fixed coverage floor is up to logarithmic factors, and the exact IUT attains it: menu richness costs times the abstention odds, even though validity costs nothing. The lower bound holds for audits that know the score distribution; the upper bound extends to any menu of VC dimension . We then explain the linear cost. Every certificate that never rejects while some rule looks feasible on the audit data—the IUT, Bonferroni, DKW and uniform-convergence bounds—pays it even at laws where samples suffice, because the best-looking rule is optimistically biased. We then identify when the gap can and cannot be avoided: against constant-risk alternatives a reference-free audit reaches the testing scale, including the square-root rate on the sign family in its stated small-gap regime, whereas over a rich class with unknown heterogeneous risk the linear cost is unavoidable up to a logarithm. In simulation with , the raw menu cost changes 7.7-fold between and ; after dividing by the abstention odds and the theorem-motivated localized-VC logarithmic normalization used in Figure 1, the two endpoint values differ by about 5%. Held-out model families show that the same feasible-mixture mechanism is not confined to synthetic constructions: across three contemporary VLM judges, all 60 split-specific confidence-gate families have positive exact-hiding windows and 28/30 contracts selected on audit-A remain hidden on untouched audit-B; on the cleaner multi-annotator Polaris target, hiding persists on all 20 splits but is much narrower and transfers in 3/10 resplits. Frozen image and multimodal classifiers separately show boundary behavior consistent with the exact IUT guarantee and strong power relative to the practical certificates we compare.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.