CURE: Counterfactual Unlearning of Large Language Models via Density Ratio Estimation
Abstract
Machine unlearning for large language models aims to remove the influence of designated training data while preserving behavior supported by the remaining data. Yet knowing what to forget does not determine what the model should predict after forgetting: the factual prediction entangles retain- and forget-supported influences, making the desired retain-only behavior unidentifiable from the original model alone. We show that this ambiguity can be resolved by introducing a controlled second observation that amplifies the forget influence. Under a retain–forget mixture model, our Bidirectional Mixture Identification result reduces recovery of the retain-only next-token distribution to estimating a context-dependent forgetting ratio. We further show that this ratio is characterized by the density ratio between the factual and forget-amplified models and can be updated autoregressively during generation. Based on these results, we propose CURE (Counterfactual Unlearning via Ratio Estimation), which uses a forget-aligned auxiliary model to reveal the direction of removal and contextual ratio estimation to determine its magnitude, while keeping the factual model fixed. We additionally analyze how model-realization, auxiliary-alignment, and ratio-estimation errors affect recovery. Across TOFU, MUSE-News, and RealToxicityPrompts, CURE achieves favorable forgetting–utility trade-offs; on TOFU-10%, it attains forget/retain ROUGE-L of 38.3/98.4, compared with 39.6/99.2 for the retain-only reference. These results support counterfactual unlearning as an identification-and-estimation problem rather than solely an objective for suppressing forget-associated behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.