Split Conformal Interval Length under Huber Contamination
Abstract
Existing analyses of interval length in split conformal prediction largely focus on clean data. However, how contamination at the training and calibration stages alters interval length relative to the corresponding clean population oracle remains understudied. In this paper, we consider Huber contamination and derive a finite-sample upper bound that holds uniformly over contaminating distributions, with leading terms scaling as , in addition to a clean-sample estimation term. We further establish a local minimax lower bound that matches an attainable fixed-confidence upper rate, up to constant factors, on a regular asymmetric location submodel. For calibration contamination, we obtain an exact worst-case characterization of the population interval-length deviation. We also extend the upper-bound results to conformalized quantile regression and adaptive residual scores. Experiments on synthetic and real-world datasets illustrate the distinct effects of training and calibration contamination and are consistent with our theoretical analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.