Active Unlearning Probing: Label Inference Attack from the Client's Vantage in Federated Unlearning
Abstract
Federated unlearning allows a client to request that its data be retracted from a jointly trained model, after which the server broadcasts the updated model to all participants. Prior forgotten-label inference typically assumes an attacker with a server vantage that can read individual client updates, yet in practice the server is usually trusted, so we ask whether the same inference attack can be mounted from the vantage of an ordinary client. Such a threat model is less restrictive and easier to realize. We show that the answer is yes. We propose Active Unlearning Probing (AUP), which issues short unlearning probes as a legitimate client to build a per-class gradient template, then matches the classifier-layer direction of the observed global-model change against these templates to infer which label class another user has forgotten. Using only its own labeled training data, with no shadow set, no labeled test set, and no queries to the victim, AUP identifies the forgotten label at 0.96 accuracy on CIFAR-10 and 0.80 on MNIST, likewise identifies labels on CIFAR-100 with a ViT-B/16 and on the Yahoo Answers text dataset, recovers three simultaneously deleted classes at an intersection-over-union (IoU) of 0.58, and remains effective under two different unlearning mechanisms. We also find that AUP becomes more effective as the federation grows, a phase transition governed by the signal-to-noise ratio rising with the number of clients. Code is available at https://anonymous.4open.science/r/AUP.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.