Understanding and Mitigating the Overcorrection in Machine Unlearning via Feature Learning
Abstract
Machine unlearning aims to remove the influence of particular data points from a pre-trained model in response to growing privacy and ethical concerns. While gradient-ascent (GA)-based unlearning provides a simple and widely used strategy, its optimization dynamics remain poorly understood. In this work, we identify a striking non-monotonic behavior in GA-based unlearning, which we call the overcorrection phenomenon: the forgetting quality evaluated only on the forgetting data first improves and then deteriorates as unlearning proceeds. Specifically, the Kolmogorov-Smirnov (KS) distance between the truth-ratio distributions of the unlearned model and the retrained model first decreases and then increases, indicating that the unlearned model initially approaches the retrained reference but later deviates from it. To explain this phenomenon, we provide a three-stage theoretical characterization under a two-layer convolutional neural network model. Our analysis specifies when the forgetting quality enters a favorable regime, how long it remains there, and how it eventually deteriorates. The theoretical results further reveal the key role of the learning rate, motivating a simple learning-rate adjustment strategy that mitigates overcorrection. Experiments on both synthetic and real-world datasets support our theoretical predictions and demonstrate the effectiveness of the proposed strategy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.