Twenty Logprobs Are Not Enough: The Feasibility Frontier for Conformal Prediction under Black-Box Output Truncation
Abstract
Conformal prediction assumes the analyst can score every candidate label. Production APIs violate it: the interfaces we surveyed return at most 5–50 log-probabilities per position, some none. We characterize what distribution-free coverage survives this truncation. Coverage is capped at top- accuracy plus a tail-concentration term we define; for rules that neither single out hidden labels nor reach past the returned window the bound reads : the interface, not the score, limits confidence. A matching truncated procedure attains the bound marginally (not given a usable set), and a binomial criterion with a phase transition at governs collapse: the population budget is necessary but – too small at practical . We verify the frontier over 670,000 scored predictions and a 49,696-call audit of seven production endpoints; two deliver less than the access level their schema advertises (one at the temperature a classifier would use, one in default reasoning mode), and returned probabilities imply censoring rates up to below the realized ones. Where the tail has an atom the term is an escape route: one calibration-identified appended label recovers 53% offline and 74–100% on six of seven live endpoints. Feasibility must be certified from labels, not from the API's numbers. Code and records are in the supplement and at https://github.com/xxx/xxx (public upon acceptance).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.