acceptodds
Under review as a conference paper at ICLR 2027

Unlearning Can Hurt When There Is Nothing to Forget

Abstract

Unlearning is meant to remove one piece of knowledge from a language model and leave the rest intact. Its side effects are measured by how much the model's scores on other topics drop, and the whole drop is charged to the removal. We show that for two widely used methods, RMU and NPO, most of that drop does not come from removing that knowledge. We fine-tune two copies of a model the same way, one on the documents to be forgotten and one on unrelated documents, and give both the identical unlearning run. If the damage came from removing what the model learned, the copy that never learned it would be spared. Under RMU it is not: across three datasets and fictitious authors, five models from four families, 14 fine-tuning seeds and four measures of damage, it loses about as much as the copy that did, including on the closely related topics where damage is largest, and unlearning statements a model never saw costs as much as unlearning what it knows from pretraining. At least 74% of RMU's collateral damage on MMLU, and over 90% on the other two datasets, occurs without the model having learned what was removed, and NPO on MMLU, at the same amount of damage, behaves like RMU. The control detects dependence where it exists: under GradDiff and SimNPO, a copy that learned other documents about the forget topic is damaged less than the taught copy, about half as much under GradDiff, and damage aimed at chosen topics, by gradient ascent or by RMU itself, is detected in nearly every draw. Unlearning evaluations should report collateral damage next to the damage the same run does to a model that never learned the forget documents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.