acceptodds
Under review as a conference paper at ICLR 2027

How Well Does LLM Unlearning Generalize?

Abstract

LLM unlearning methods aim to remove a set of knowledge (the forget set) from a trained language model while keeping another set (the retain set) intact. Most existing unlearning benchmarks and methods construct both sets in a single form—either question–answer pairs or verbatim text continuations—but knowledge removed in one form may remain recoverable in many others in deployment, such as a paraphrase in the same form, the other form, another language, or a promptbased attack. In this paper we study how unlearning generalizes across query forms and languages, and how to improve the generalization through synthetic data augmentation and localized training, over fourteen unlearning objectives, two model families, eight query forms, four prompt-based attacks and up to eighteen languages. We report three findings. First, unlearning does not generalize well across the question–answer form and the verbatim form, which a model stores largely independently. Second, unlearning generalizes across most languages, with the exception of low-resource languages that spend more layers on language-specific processing. Third, among individual modules, the unlearning weight update generalizes best in the feed-forward layers, while the value and output projections give a surface-level unlearning that does not generalize and the query and key projections unlearn the least, in both standard transformers and linear-attention hybrid models. Building on these findings, we release AUGUNLEARNING, a suite that improves the generalization of both the training and the evaluation of existing unlearning methods with a synthetic data augmentation and a new metric. Together, our findings expose the vulnerabilities of current unlearning methods and the blind spots of current benchmarks. Our proposed methods solve the cross- form generalization problem and improve existing unlearning methods towards a better trade-off between unlearning quality and model utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.