Bootstrap-Zoom: Inference after Checkpoint Selection
Abstract
Validation data are often used both to select a checkpoint and to estimate its population risk. This reuse can invalidate ordinary confidence intervals, especially when many correlated checkpoints compete. Bootstrap-Zoom extends the zoom correction by replacing its known joint tail with a studentized multiplier-bootstrap tail estimated from the validation loss matrix used for selection. Because the tail, the variance scales, the observed gaps, and the selected checkpoint all depend on one sample, approximation guarantees for fixed regions do not by themselves justify the resulting confidence set. Conditional on a training path fixed before validation, we bound the probability of abstaining or covering from below, accounting for this dependence, finitely many bootstrap draws, and grid discretization. When population standard deviations are bounded away from zero and the checkpoint count and grid satisfy growth conditions, we derive coverage-deficit rates at both positive and zero variance floors. An exact endpoint search returns the same interval as testing every grid value, without assuming that the accepted set is connected. On seven combinations of training setting and loss, we separate the width saved by checkpoint dependence from the width saved by gap discounts. Selected-risk overcoverage can coexist with joint undercoverage, and a rare-loss counterexample shows that positive empirical variances alone do not justify Gaussian calibration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.