acceptodds
Under review as a conference paper at ICLR 2027

SELECTION-SAFE CONFORMAL EARLY EXIT: DISTRIBUTION-FREE COVERAGE UNDER ADAPTIVE DEPTH

Abstract

Calibrating a conformal prediction set at each early exit does not preserve coverage after a data-dependent exit is selected. We construct one terminal path score from prefix maxima, calibrate one split-conformal quantile, and return nested prefix sets that all contain the valid terminal set. The resulting guarantee is selection-safe without an α/L correction and holds simultaneously for every post-calibration label-free policy that preserves terminal-set containment. We further bound the efficiency cost of the path maximum by the empirical tail of the early-over-terminal score excess, linking set inflation to the mechanism diagnostic J(x, y). At the smallest admissible excess threshold, the held-out plug-in right- hand side is 0.016–0.086 against measured mean positive terminal-set inflation 0.01–0.08; the cell-wise ratio is 1.07–1.73. Across eight matched encoder task– backbone cells, including an RTE split with ncal = 498 drawn from the training partition, pathwise selection attains macro coverage 0.901 versus 0.814 for naive layerwise selection, with average set size 1.41 and latency fraction 0.71. On CIFAR-100 with a 24-exit ViT-L/16 and ncal = 500, first-eligible stopping at budget b = 4 attains coverage 0.902, average set size 2.76, and latency fraction 0.57; a design-chosen restriction reduces the returned set to 2.52 while, necessarily, increasing latency to 0.61. The same terminal calibration supports deployment-time budget changes without refitting. A joint L–n sweep isolates candidate-count scaling. At n = 100, pathwise set size changes from 3.02 to 3.11 as L increases from 4 to 24, whereas Bonferroni becomes exactly degenerate at L ≥ 12 and returns all 100 classes. Against MinSE valid selection, the paired set-size difference at (n, L) = (100, 24) is 0.80 [0.47, 1.13], so the long-path, limited-calibration advantage is assessed from paired replicates rather than overlapping marginal intervals. Independent-depth CIFAR-100 and ImageNet-1k replications preserve the qualitative ordering. Full-vocabulary decoder experiments remain a negative boundary, with average sets containing 7,162–9,620 tokens and nearly all examples reaching full depth. The empirical picture is therefore specific: pathwise calibration is most useful when the number of candidate exits is large relative to calibration size, when deployment budgets change after calibration, and when shallow heads can return actionable sets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.