acceptodds
Under review as a conference paper at ICLR 2027

Early Memories Don't Fade Easily: On the Difficulty of Unlearning Pretraining Knowledge

Abstract

Machine unlearning aims to remove the influence of specific training data from a deployed model. Existing benchmarks introduce the forget set through a brief fine-tuning stage after pretraining, so they primarily evaluate the removal of recently acquired knowledge. In practice, however, deletion requests may concern data encountered during pretraining itself. Whether these two settings are equally difficult remains poorly understood. We study how the stage at which knowledge is acquired affects its subsequent removal. We hypothesize that data encountered earlier in training can become more strongly integrated with representations that are subsequently reused, whereas data introduced later may induce more localized changes. We theoretically analyze this training-history effect, first characterizing how forget-set influence accumulates along the optimization trajectory and then showing, in a two-layer linear model, that earlier introduction of the forget data leads to greater parameter displacement from the retain-only solution through shared feature directions. To test this hypothesis, we train language models from scratch while controlling when the forget set is introduced. Across different models, unlearning methods, and evaluation metrics, we find that knowledge maintained from earlier stages of pretraining is consistently more difficult to remove than the same knowledge introduced later. These results show that two models can fit the same forget data similarly yet differ substantially in how that knowledge is represented in their parameters and how readily it can subsequently be removed, depending on when the data entered training. Our findings motivate benchmarks that explicitly account for training history and the development of unlearning methods that are effective across more general and realistic training settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.