Analyzing LLM Unlearning: Token-Triggered Forgetting Mechanism and Entity-Weighted Remedy for Trigger-Free Queries
Abstract
Large language model (LLM) unlearning aims to suppress targeted training knowledge while preserving unrelated capabilities. We identify a *token-triggered forgetting* mechanism associated with favorable forget–retain trade-offs in several gradient-based objectives and low-rank updates: forget-target tokens attract attention, accompanied by selective collapse of the forget representation subspace. Controlled trigger injection and activation patching provide functional evidence for this pathway. The mechanism is vulnerable on *trigger-free* queries, where the target knowledge is inferred rather than explicitly named. We demonstrate this failure on TOFU and reproduce its pattern on RWKU. To mitigate it, we propose EE-TNPO, which uses early-exit distributions to up-weight knowledge-bearing answer tokens during unlearning. By emphasizing these tokens, which include names and factual attributes, EE-TNPO reduces trigger-free leakage while preserving general utility on TOFU and RWKU. Moreover, WMDP and MUSE results support its application to corpus-level unlearning without entity-defined forget targets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.