Retrained Models Can Relearn in Context
Abstract
Machine unlearning aims to make a language model behave as if it had never been trained on certain data, and audits check whether it does. A common audit, in-context relearning, counts a fact as not forgotten if related context brings the answer back. A model retrained without the fact fails this audit too when the fact can be inferred from what the model kept. We build synthetic facts in which a person's sector follows from their employer and ask models to forget each person's sector: in seven combinations of Llama and Qwen models and fact sets, a model retrained without the sectors selects the forgotten sector from ten candidates 54–91% of the time once told the employer, and retrained OLMo models do the same. In real pretraining, a model from the Hubble suite never shown a set of biographies infers a person's nationality from their birthplace as often as the model that saw them (98%). Recovery alone therefore does not show that unlearning failed; what matters is whether the unlearned model behaves differently from a retrained one. Judged this way across 145 hyperparameter settings of 11 unlearning methods, open-ended generation fails to detect most unlearned models that still prefer the forgotten answer, and a score based on how much context raises the model's preference for the answer rates 45 of 48 as more forgotten than the retrained model; a multiple-choice question without context detects most of them. For facts that can be inferred, we add conflicting context: a prompt that names a wrong employer. A retrained model follows it; an unlearned model that keeps the original answer departs from retraining. This test flags RMU models on Qwen3 and Llama that pass the multiple-choice question, provided they still infer the sector from the employer. For facts that cannot be inferred, such as random identifiers inserted into Hubble's pretraining data, recovery from context remains a valid test: seven of nine 8B models unlearned with SimNPO or WGA that pass our health checks choose the forgotten identifier from 11 options 89–98% of the time, against 11% for a model pretrained without it. Unlearning audits should run a retrained model through the same prompts and, for facts that can be inferred, add conflicting context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.