Unlearning Can Remember the Test and Forget the Rest
Abstract
Unlearning should remove specific facts from a model and leave everything else intact. Evaluations check the second part with a retain set: facts the model should still know afterwards. But many unlearning methods train on that same retain set. A method could therefore damage much of what the model knows, restore exactly the facts that will be tested, and look nearly as good as retraining from scratch. We show that this happens. We train language models of 0.6B to 1.24B parameters from two families on synthetic biographies, so we can also test people the model knows but no method targets or trains on. A method that resets parameters and then fine-tunes on the retain set can remove the targets while keeping retain accuracy high, yet in all 17 configurations of our main setting these untouched people lose 44 to 104% as much accuracy as the targets; where we retrained a matched model without the targets, the untouched people lose only 6 to 15% as much. The loss appears in both model families and equally for untouched facts whose questions were never trained. Prompting with their own biographies does not bring most of their facts back. The loss persists even when the untouched people are learned exactly like the targets, and only people in the retain set are spared. On the TOFU benchmark, a standard method, gradient difference with a strong retain weight, shows the same pattern: retain questions it trains on stay at the retrained level, while untouched questions of the same kind fall 20 points below it. A retain score therefore cannot show that unlearning was selective, and evaluations should also test known facts that no method trains on.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.