ARC-Unlearn: Blast-Radius-Aware LLM Unlearning via Agentic Retain-Set Synthesis and Adjacent-Retain NPO
Abstract
Machine unlearning for large language models aims to remove unwanted target knowledge while preserving the related knowledge around it. However, most existing approaches primarily evaluate how effectively target information is forgotten, potentially overlooking damage to semantically related but legitimate knowledge. We refer to this localized collateral damage as the blast radius. We introduce BlastBench, an adjacency-aware benchmark for quantifying collateral damage during machine unlearning through Adjacent Retain Rate (ARR). We further present ARC-Unlearn, a multi-agent framework that identifies target-specific adjacent concepts, constructs an adjacent-retain set, and protects this knowledge during unlearning. Its optimization component, Adjacent-Retain NPO (AR-NPO), combines NPO-based forgetting, adjacent-retain cross-entropy, and general-knowledge KL regularization. Across three 7B-scale models and 45 targets, standard forgetting baselines achieve high forget efficacy but substantially degrade adjacent knowledge. Among the evaluated methods, AR-NPO is the only one that achieves non-trivial forget efficacy () while improving adjacent retention relative to the original model (). An -sweep reveals a controllable FE–ARR trade-off between target forgetting and adjacent retention. On a matched 9-target subset at 32B scale, AR-NPO reaches and , with all nine targets satisfying our surgical-zone criteria. Full-parameter fine-tuning further achieves near-complete measured forgetting () together with above-baseline adjacent retention (). Finally, held-out and cross-generator audits show that the adjacent-retention gains persist beyond the original generated QA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.