From Per-Class Summaries to Top- Decisions: Sharp Information Limits
Abstract
When deploying a classifier ahead of an expensive reranker or human expert, practitioners must choose a shortlist size that captures the true label while limiting downstream costs. Can this shortlist size be reliably chosen from standard evaluation reports containing top-1 accuracy and per-class score distributions? We show that these reports need not suffice: two evaluation sets can share identical top-1 accuracy and per-class summaries, yet yield substantially different top- coverages. This ambiguity arises because per-class summaries discard how scores across different classes co-occur on individual examples. To understand the fundamental limits of this information loss, we first analyze the idealized setting of true posterior probabilities. Even with ideal scores, the maximum coverage gap between report-compatible distributions is at top-2 and at top-3, converging to as both shortlist size and class count grow. We establish these sharp limits using dual certificates and matching constructive examples, which provide both optimistic and guaranteed shortlist bounds for any supplied report. For general uncalibrated classifiers, exact programs certify actual coverage and shortlist requirements. Across 150 benchmark reports, required shortlist sizes are uniquely determined at 90% target coverage, but 29 reports remain ambiguous at 95%—in one case swinging the required shortlist from two to five candidates. Adding joint cross-class statistics narrows this uncertainty, while a compact true-label rank histogram eliminates it entirely. Our results pinpoint when standard per-class summaries suffice for shortlist sizing and what additional statistics evaluators should report.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.