acceptodds
Under review as a conference paper at ICLR 2027

A Certified Robustness Framework for Backdoor Poisoning Attacks using -Divergences

Abstract

Machine learning systems trained on untrusted data are vulnerable to data and backdoor poisoning attacks, in which an adversary corrupts training and test samples to manipulate predictions at inference time. Randomized smoothing yields provable defenses against such attacks, but existing applications to train-time robustness are limited to specific threat models and smoothing mechanisms. We propose a general framework that allows expressing complex adversary constraints on the number of corrupted training examples, feature and label perturbation magnitudes, and test-time corruptions. Our framework randomizes inputs by combining data subsampling with feature, label and test-sample noising. We propose a composable certification procedure based on -divergences that handles arbitrary combinations of threat models and noising strategies, and recovers several prior certification regimes exactly as special cases. We demonstrate our framework experimentally by certifying, for the first time, general -constraint poisoning adversaries under bagging with additive Laplace, Gaussian, and Uniform noises.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.