Forget Narrowly, Retain Broadly: Unlearning as an Asymmetric Generalization Problem
Abstract
Machine unlearning in LLMs is the targeted removal of specific knowledge while preserving all other capabilities. Existing benchmarks measure it unreliably. They miss knowledge that resurfaces under paraphrased or indirect queries, a failure we call *under-forgetting*, and lack the semantic, syntactic, and lexical probes needed to verify that unrelated knowledge is preserved, a failure we call *over-forgetting*. Both failures reflect an asymmetric generalization problem. Forget evaluation must cover diverse query formulations of the same target facts, testing whether forgetting holds beyond training prompts. Retain evaluation must probe a far larger and implicitly defined set: every fact disjoint from the forget target. The retain set thus defines the effective forget set, yet current datasets provide no fine-grained annotation of this forget-retain boundary. We address this with SUITE, an evaluation protocol and training corpus that captures forget-retain structure for real-world factual domains. SUITE improves *all* unlearning methods substantially, showing that training data is as important as algorithmic design. Moreover, we introduce JensUn++, an unlearning algorithm that achieves a better or equal forget-retain utility trade-off across three LLMs, in both sequential and joint unlearning settings. JensUn++ is the only method producing no gibberish answers when asked about forget facts. Code and datasets are available at [this anonymous link](https://anonymous.4open.science/r/forget-narrowly-retain-broadly).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.