acceptodds
Under review as a conference paper at ICLR 2027

The Iterated-Logarithm Price of Post-Hoc Threshold Selection

Abstract

The standard way to deploy a selective predictor is to calibrate it, look at the risk coverage curve, pick the operating point that looks best, and report the risk certificate at that point. Every widely used certificate such as Clopper–Pearson, Hoeffding, empirical Bernstein, betting bounds, and Learn-then-Test is valid only at a threshold fixed before the calibration data are seen, so none carries a uniform guarantee under it. Some happen to survive specific attacks at specific scales anyway, but by accident rather than by design. We give a certificate built for the workflow directly. Sorting items by a confidence score is a function of the covariates alone, so it leaves the losses conditionally independent; the resulting sorted filtration turns the risk coverage curve into a martingale problem, and anytime valid concentration applies directly. This yields a band Uk that covers the realized risk of every prefix simultaneously, for an arbitrary score, with no assumption that the score is calibrated or even informative. And it does so, we show, with the certificate paying no extra price for dependence within blocks specified in advance. The band’s√︁ ln ln k/k width is not slack: any band valid at every prefix simultaneously must be at least this wide infinitely often, and we check the rate empirically across calibration sizes spanning a factor of 25. Across three language models, seven score families, five datasets and a vision model, the band was violated 0 times in 32,300 replications with exactly known conditional risk and 0 times in 320 genuine decoding replicates. In a deployment protocol that chooses the threshold after seeing the calibration curve, Clopper–Pearson and a betting certificate are exceeded on held-out data 12.7% and 11.6% of the time while ours is exceeded 0 times in 24,000. We trace the betting certificate’s failure to an edge case in our own implementation a 32% silent-margin defect under heterogeneous risk against 1% on i.i.d. data and repair it with a one-line, provably valid patch that resolves the failure under a fixed threshold and substantially, though not completely, under post-hoc selection. A feasibility condition states exactly when a population-level version of the same certificate can reach a target risk level, and we report the regime, small calibration sets and mid-accuracy models, where it cannot, along with how much a lower-error subset closes the gap. Code is available at https://anonymous.4open.science/r/iterlog-risk-E131/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.