acceptodds
Under review as a conference paper at ICLR 2027

When Deletion Writes: Target Imprints in LLM Editing and Unlearning

Abstract

Machine unlearning and model editing aim to remove specific knowledge from language models while preserving unrelated knowledge. A common approach to evaluating deletion is to test whether the deleted knowledge can still be recovered, where successful recovery is interpreted as evidence that the knowledge has survived. In this work, we show that such an interpretation can be misleading, because the deletion process itself can create a recoverable trace of the target it suppresses. We demonstrate this by applying the same deletion procedure to random target tokens that the model has never associated with the subject, and we find that these tokens become recoverable after deletion. Since these tokens had no prior association with the subject, the signals that newly emerge for them cannot represent surviving knowledge. In other words, a deletion can write the very thing that it deletes. We refer to this phenomenon as a loss-target imprint. Across the mechanisms we study, the imprint appears when the deletion loss directly references the target token, increases with deletion strength, and remains detectable by probes adapted to the deletion method, with the target often ranking first or second among all vocabulary tokens. These findings expose a fundamental limitation of recovery-based evaluation: failure to recover a deleted target does not establish that the target has been removed, whereas successful recovery does not establish that the recovered signal survived deletion. Our results highlight the need for unlearning evaluations that distinguish knowledge surviving deletion from signals introduced by the deletion process itself.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.