acceptodds
Under review as a conference paper at ICLR 2027

Sample-Wise Early Stopping for LLM Unlearning

Abstract

Large language model (LLM) unlearning aims to remove the influence of designated training data while preserving model utility, with applications to privacy protection and regulatory compliance. However, existing methods typically use a fixed training budget, overlooking differences in unlearning difficulty across samples. Consequently, some samples remain insufficiently unlearned, whereas others undergo excessive unlearning, resulting in redundant computation or degraded model utility. To address this challenge, we propose Sample-Wise Early Stopping (SWES), an efficient framework for adaptive sample-wise LLM unlearning. SWES measures each sample’s unlearning progress by comparing its length-normalized negative log-likelihood before and during unlearning. It combines temporal smoothing with threshold-based control to dynamically suspend or resume each sample’s participation in backpropagation, reducing unnecessary updates and focusing subsequent training on samples that have yet to reach the unlearning target. As a plug-and-play framework, SWES can be integrated into a range of gradient-based unlearning methods. Experiments on TOFU and MUSE demonstrate that SWES improves the trade-off between unlearning effectiveness and model utility in multiple settings, while reducing cumulative sample participation in backpropagation by 15.9%–64.6% and achieving training speedups of 1.15×–2.34×.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.