acceptodds
Under review as a conference paper at ICLR 2027

Relearning Resistance Under Adam: A Gradient-Entropy Certificate for Machine Unlearning

Abstract

Relearning provides a direct test of whether machine unlearning has removed information or merely suppressed its immediate behavior. Relearning Convergence Delay () quantifies this resistance by accumulating excess forget-set error along the relearning trajectory. Existing theory focuses on gradient descent, despite the widespread use of adaptive optimizers such as Adam for training modern models. We develop a framework for analyzing under Adam and derive an upper bound involving three quantities at the unlearning initialization: the forget-set loss gap, the local Hessian condition number, and Adam's second-moment condition number. Because large models are often fine-tuned by updating only a restricted subset of parameters, we further show that the expected certificate can be bounded in terms of the entropy and scale of the forget gradient. This analysis motivates Gradient-Entropy Maximization (GEM), a baseline-agnostic regularizer that promotes both gradient dispersion and magnitude. We establish an convergence rate to an -stationary point for the GEM objective and evaluate it empirically. Across classification, diffusion, and language-model unlearning, GEM can improve resistance to Adam relearning while largely preserving retain performance and forgetting quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.