acceptodds
Under review as a conference paper at ICLR 2027

Provably Robust Conformal Prediction against Arbitrary Data Poisoning Attacks

Abstract

Conformal Prediction (CP) provides a principled approach to uncertainty quantification in machine learning by constructing prediction sets that provably contain the true label with high coverage. However, recent studies have shown that CP is highly vulnerable to data poisoning attacks: manipulating even a small fraction of training or calibration samples can significantly inflate prediction set sizes or violate coverage guarantees, undermining CP’s reliability. We propose JointCP, a robust CP framework to defend against arbitrary data poisoning attacks with formal coverage guarantees, i.e., the coverage is still valid when a fraction of training and calibration samples used in CP are arbitrarily perturbed (e.g., deleted, added, modified, or combined in any manner). Results on multiple image classification datasets demonstrate the effectiveness and robustness of JointCP from both empirical and theoretical perspectives.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.