Advancing Machine Unlearning Evaluation Requires Rethinking Retraining
Abstract
Machine unlearning is emerging as a pivotal paradigm aimed at reducing the impact of unwanted data points previously used in training machine learning models. Despite recent advancements in developing new machine unlearning algorithms, robustly evaluating the effectiveness of these algorithms remains a challenging problem. Retraining a model from scratch, which many algorithms seek to circumvent, continues to be a popular benchmark for comparison due to its inherent nature of excluding undesirable data from the outset. In this study, we critically assess the role of retraining as a benchmark for evaluating machine unlearning algorithms. Through extensive experiments across various evaluation metrics, datasets, and model architectures, we uncover notable instability in retrained models. Our findings reveal that retrained models can exhibit both the highest and lowest effectiveness in unlearning, while maintaining the same utility. Additionally, this phenomenon has been shown to be consistent across multiple membership inference attack and distance function metrics. Although our results do not negate the importance of retraining as an unlearning method with guaranteed data removal, they highlight the need for its careful application in machine unlearning evaluation. Code for reproducing all results will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.