acceptodds
Under review as a conference paper at ICLR 2027

Semantic LLM Unlearning through Causal-Inspired Self-Distillation

Abstract

Machine unlearning is developed to remove sensitive and hazardous information that Large Language Models (LLMs) have inadvertently memorized. Most existing LLM unlearning methods operate at the lexical level, suppressing only a specific subset of tokens related to the underlying semantic concepts. Consequently, forgotten knowledge remains accessible via paraphrasing or temperature sampling, so these methods fail to achieve semantic unlearning. To address this issue, we take a causal-inspired view of unlearning as a concept-level intervention that disrupts the dependency between sensitive information and internal model representations. Guided by this view, we propose the Semantic-aware Path Intervention Distillation (SPID) framework. SPID utilizes self-distillation optimized via reverse-KL divergence to suppress the entire semantic cluster of the forget concept. Furthermore, to quantify the degree of semantic unlearning, we introduce Semantic Residual Knowledge (SRK), which detects semantic leakage by measuring the model's belief regarding the forgotten fact. Our experiments demonstrate that SPID achieves superior knowledge erasure with minimal impact on model utility, while SRK provides a more faithful evaluation of semantic unlearning than traditional metrics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.