acceptodds
Under review as a conference paper at ICLR 2027

When Unlearning Is Not the End: Controlled Oracle Recovery Against Federated Unlearning

Abstract

Federated Unlearning (FU) aims to remove the influence of client data on a shared large language model upon request. However, when the target model remains trainable, subsequent updates may restore the behavior intended to be unlearned. This creates an integrity threat: a malicious participating client may counter another client's deletion request by steering the shared model back toward its pre-unlearning behavior. Existing evaluations largely focus on the immediate post-unlearning state and largely overlook the possibility of such a recovery case. We introduce Controlled Oracle Recovery, an anti-unlearning audit that evaluates whether a deletion event reveals its target records and whether small target-guided updates can reverse forgetting more effectively than compute-matched control updates. Deletion target inference reaches at least \(0.999\) ROC AUC in 10 of 24 evaluations. After 100 local update steps, recovery using the deletion targets exceeds the strongest matched control in all 48 comparisons. In multiple settings, recovery also restores forgotten-answer probability to its pre-unlearning level while retaining at least \(95%\) of the utility measured before unlearning. The results reveal a broad but heterogeneous anti-unlearning susceptibility and establish a necessary local condition for a FU rollback attack on LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.