acceptodds
Under review as a conference paper at ICLR 2027

When the Bubble Bursts: Malicious Unlearning Requests Undermine Federated Learning

Abstract

Growing demand for the “right to be forgotten” has motivated federated unlearning (FU), which aims to efficiently remove the influence of designated training data while preserving model utility. However, we identify a security gap in existing FU approaches: treating client-issued unlearning requests as trusted inputs allows adversaries to exploit the unlearning service itself to manipulate model predictions. To expose this vulnerability, we propose BURST, an attack framework that induces misclassification of selected samples through malicious unlearning requests. Specifically, the adversary constructs a target-support set and an interference set from public auxiliary data and incorporates both into federated training, using these external samples to exert controllable adversarial influence on the target prediction. The support set reinforces the model’s correct prediction for the target sample, while the interference set is designed to undermine it. During the unlearning phase, the adversary submits malicious requests to unlearn the support set and strengthens the interference set’s adversarial influence, jointly turning previously correct predictions into misclassifications. Extensive experiments show that BURST achieves an attack success rate of 91.43% on CIFAR-10 while reducing global model accuracy by only 5.66%. Among misclassified targets, confidence in the predicted labels is on average 30.74% higher than confidence in the ground-truth labels. Our findings demonstrate how client-controlled unlearning requests and data can jointly compromise FU, highlighting the need to validate both to safeguard model integrity. Code is available at https://anonymous.4open.science/r/Code-for-BURST-4CC1.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.