Semantic LLM Unlearning through Causal-Inspired Self-Distillation
Abstract
Machine unlearning is developed to remove sensitive and hazardous information that Large Language Models (LLMs) have inadvertently memorized. Most existing LLM unlearning methods operate at the lexical level, suppressing only a specific subset of tokens related to the underlying semantic concepts. Consequently, forgotten knowledge remains accessible via paraphrasing or temperature sampling, so these methods fail to achieve semantic unlearning. To address this issue, we take a causal-inspired view of unlearning as a concept-level intervention that disrupts the dependency between sensitive information and internal model representations. Guided by this view, we propose the Semantic-aware Path Intervention Distillation (SPID) framework. SPID utilizes self-distillation optimized via reverse-KL divergence to suppress the entire semantic cluster of the forget concept. Furthermore, to quantify the degree of semantic unlearning, we introduce Semantic Residual Knowledge (SRK), which detects semantic leakage by measuring the model's belief regarding the forgotten fact. Our experiments demonstrate that SPID achieves superior knowledge erasure with minimal impact on model utility, while SRK provides a more faithful evaluation of semantic unlearning than traditional metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.