acceptodds
Under review as a conference paper at ICLR 2027

UNISON: Joint Privacy and Backdoor Forgetting for Full-Information Deletion of Poisoned Training Data

Abstract

Machine unlearning aims to remove the influence of selected training examples without retraining from scratch. We study deletion when the examples to be forgotten also taught the model malicious behavior. Backdoor poisoned samples can leave a sample-level privacy trace and a trigger-dependent rule that generalizes beyond the samples themselves. Our experiments show that existing methods do not remove these effects together. Privacy-focused unlearning can reduce membership leakage while the backdoor remains active. Backdoor repair can reduce attack success while membership information remains detectable. Increasing either intervention eventually damages clean accuracy without reliably solving the other objective. We trace this separation to the forgetting gradients. Their overall alignment is low, but a small set of directions repeatedly supports both objectives across batches. We introduce UNISON, which identifies this stable shared structure, preserves the remaining objective-specific directions, suppresses updates that conflict with retained knowledge, and shifts effort toward the objective with the larger residual gap. Across CIFAR-10, CIFAR-100, Tiny ImageNet, and ImageNet-1K with four backdoor attacks and four image classifiers, UNISON reduces mean attack success rate to approximately 5.2% and membership-inference deviation to approximately 0.029, while remaining within 0.9 percentage points of retraining accuracy. These results show that poisoned data deletion requires balancing distinct forms of forgetting rather than optimizing privacy or behavior alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.