A Ground-Truth Model Suite for Unlearning Evaluation Across the LM Training Lifecycle
Abstract
Machine unlearning – the ability to selectively remove specific information from a trained model – is critical for privacy, copyright compliance, and safety. In this work, we present a controlled study of machine unlearning across the large language model (LLM) development lifecycle, spanning pre-, mid-, and post-training. To achieve this, we inject multiple distinct datasets during the from-scratch training of OLMo-2 models ranging from 179M to 2.7B parameters and trained on 210B tokens. These datasets encompass seven diverse unlearning tasks with varying difficulty levels, including the removal of Gaussian poisons, fictional knowledge, reasoning patterns, social interaction data, and adversarial attacks. Crucially, we also provide counterfactual models trained without these datasets, which serve as exact ground-truth targets for unlearning evaluation. Leveraging this suite, we investigate the effectiveness of existing unlearning algorithms across three distinct lifecycle settings: (1) immediate unlearning after pre-training, (2) unlearning after decaying from the pre-trained model and (3) unlearning via mid-training. Our experiments reveal that unlearning success is highly non-uniform where effectiveness depends heavily on the stage of the training cycle and is significantly influenced by model scale and task complexity. Surprisingly, we find that one simple and hyperparameter free baseline among eight unlearning methods is exceptionally difficult to beat: continual pre-training on fresh data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.