UnlearnBench: A Clean and Controllable Benchmark for Fine-Grained and Compositional LLM Unlearning
Abstract
LLM unlearning is essential for ensuring technical and responsible AI use, especially in addressing privacy leakage and evolving regulatory requirements. Existing unlearning benchmarks primarily evaluate unlearning through shallow question answering suppression, while overlooking several fundamental challenges: Knowledge in LLMs is acquired through complex memorization processes. It is further organized in compositional, multi-stage, and structurally entangled forms. As a result, it remains unclear whether current unlearning methods can reliably remove knowledge learned through staged training pipelines, relational reasoning, and fine-grained attribute associations. In this work, we introduce UnlearnBench, a benchmark for studying unlearning across diverse settings including single-hop QA, two-stage memorization, multi-hop reasoning, and fine-grained attribute-level forgetting. Compared with existing benchmarks, UnlearnBench better captures realistic memorization and retrieval behaviors while enabling systematic analysis of unlearning performance under clean and controlled settings. We additionally provide clean pretrained models and exact retraining checkpoints as gold-standard references for future evaluation. Extensive experiments on popular unlearning methods uncover several consistent failure modes. First, entity-level unlearning often fails even in settings where retraining successfully removes the target knowledge. Second, reasoning abilities can recover supposedly forgotten information from retained representations, revealing strong interactions between memorization and compositional reasoning. Third, attribute-level knowledge is highly entangled across entities and relation types, making selective forgetting substantially more difficult than existing benchmarks suggest. Our results demonstrate that current evaluations significantly underestimate the challenges of LLM unlearning and establish UnlearnBench as a structured testbed for future unlearning research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.