TrigCon: Certified Defenses for Conformal Prediction against Backdoor Attacks
Abstract
Conformal prediction (CP) is a statistical framework for quantifying the uncertainty of machine learning predictions, which can construct prediction sets with a user-specified probability of containing the true label. However, we observe that CP is vulnerable when the underlying classifier is backdoored, which can substantially reduce the coverage of the true label (e.g., ↓86%) while increasing the attacker-chosen target label's coverage (↑81%). Such vulnerability hinders the deployment of CP in safety-critical applications, yet still remains under-explored. Applying conventional defense approaches to backdoored classifiers poses several challenges: (1) most methods rely on accurately detecting or removing the backdoor, which may fail due to the stealthy and complex nature of backdoor behaviors; (2) these methods often target specific types of triggers and are difficult to generalize to unseen trigger patterns. In this paper, we propose TrigCon, a certified defense against backdoor attacks on CP that can provably resist arbitrary backdoors under bounded trigger sizes. Specifically, TrigCon partitions an input into several disjoint sub-inputs to construct a group of prediction sets, and then aggregates them through voting to produce the final result. This construction bounds the number of prediction sets affected by a bounded trigger, enabling us to establish a marginal coverage guarantee and derive separate trigger-size bounds for retaining the true label and excluding the target label. Extensive experiments on four benchmark datasets show that TrigCon can effectively resist backdoor attacks, achieving a high true-label coverage (e.g., 95% on Tiny-ImageNet) and a low target-label coverage (e.g., 8% on Tiny-ImageNet).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.