How Many More Labels Are Needed? Minimax Evaluation under Selective Labels
Abstract
Good observed accuracy need not establish acceptable prediction risk when difficult cases lack labels. How much can an additional audit establish? For a frozen cohort with outcome-dependent missingness, we characterize the exact fixed-query asymptotic minimax power over adaptive audits using the full observed record. A standard residual rule attains the truthful frontier with finite-sample false-alarm control. A second result identifies which risk guarantees preserve information in existing observations: count-tail and cumulative predictable-risk guarantees have the same truthful frontier but different noisy critical experiments. The latter reduce auditing to query-only testing, even with independent heterogeneous errors; the former can retain native information. Above criticality, we derive the exact robust frontier with uncertain conditional label reliability, including its blind region, and invert it to quantify the query cost of label-quality uncertainty. This separates three resources: query count, knowledge of label quality, and cohort size. Under count-tail risk bounds and truthful labels, finite certificates show that at 10% versus 20% risk and 5% false alarms, 57 queries suffice and are necessary for 95% power among 1,000 predictions with 500 pending labels. At 12% risk, even complete labels require at least 2,645 predictions for that power. These are scoped design examples, not a general finite-budget formula or validation of deployment assumptions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.