A Dynamic Example Selection Framework for Compute-Efficient Machine Unlearning
Abstract
Machine unlearning aims to remove designated training data's influence while preserving utility at a lower cost than retraining. In sample-level unlearning, forget and retain examples can share classes and features, making their selection important for balancing forgetting and preservation. Repeated full-data updates incur substantial cost. We ask which examples should receive computation to achieve an existing unlearning method's performance at a lower cost. We observe that the initial low-confidence forget subset is enriched for examples with low true-label confidence after retraining on the retain set. Motivated by this observation, we introduce Dynamic Source Selection (DSS), which refreshes retain and forget subsets using the current model's true-label confidence while preserving the base loss and update rule. DSS+ reduces rescoring by keeping selected examples in training. A Gaussian logistic analysis identifies conditions under which updates on these sources improve prediction agreement with the retrained model while preserving retain performance. Empirically, dynamic selection allocates more training to these sources. Experiments on CIFAR-10 and CIFAR-100 demonstrate that DSS can attain accuracy comparable to full-data updates at lower computational cost across multiple methods, with TOFU experiments extending applicability to language models. These results support dynamic source selection as an effective strategy for efficient unlearning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.