Early Memories Don't Fade Easily: On the Difficulty of Unlearning Pretraining Knowledge
Abstract
Machine unlearning aims to remove the influence of specific training data from a deployed model. Existing benchmarks introduce the forget set through a brief fine-tuning stage after pretraining, so they primarily evaluate the removal of recently acquired knowledge. In practice, however, deletion requests may concern data encountered during pretraining itself. Whether these two settings are equally difficult remains poorly understood. We study how the stage at which knowledge is acquired affects its subsequent removal. We hypothesize that data encountered earlier in training can become more strongly integrated with representations that are subsequently reused, whereas data introduced later may induce more localized changes. We theoretically analyze this training-history effect, first characterizing how forget-set influence accumulates along the optimization trajectory and then showing, in a two-layer linear model, that earlier introduction of the forget data leads to greater parameter displacement from the retain-only solution through shared feature directions. To test this hypothesis, we train language models from scratch while controlling when the forget set is introduced. Across different models, unlearning methods, and evaluation metrics, we find that knowledge maintained from earlier stages of pretraining is consistently more difficult to remove than the same knowledge introduced later. These results show that two models can fit the same forget data similarly yet differ substantially in how that knowledge is represented in their parameters and how readily it can subsequently be removed, depending on when the data entered training. Our findings motivate benchmarks that explicitly account for training history and the development of unlearning methods that are effective across more general and realistic training settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.