Auditing of Unlearning Algorithms
Abstract
Evaluating whether an unlearning algorithm has truly removed the influence of its forget set remains an open challenge. We introduce a hypothesis-testing auditor that uses membership-inference attacks to compute high-confidence empirical lower bounds on the zero-concentrated differential privacy (zCDP) parameter of certified unlearning, with a direct audit of the Gaussian differential privacy parameter as well; rejecting a candidate establishes that a claimed certificate cannot hold. Our bounds are computed under a black-box threat model in which the adversary additionally observes the initialization and the per-epoch shuffling order. On CIFAR-100, certified methods such as model clipping and rewind-to-delete yield small lower bounds consistent with their stated guarantees, whereas uncertified methods—including Hessian-based unlearning, interleaved descent–ascent, ascent on the forget set, fine-tuning on the retain set, and class-based forgetting methods such as BadTeacher and DELETE—often yield large lower bounds, revealing substantial residual influence of the forget set. Hessian-based unlearning is certified only for convex losses, and that certification does not transfer to the nonconvex networks we study. We further audit uncertified methods on Shakespeare and scale our evaluation to Llama-3.2-1B on TOFU, a dataset of fictitious authors. Overall, our auditor provides a practical tool for refuting unlearning claims. across model architectures, datasets, and scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.