Suppression or Relocation? Safe-Anchored Relocation for LLM Unlearning
Abstract
Large language model unlearning seeks to remove unwanted knowledge without costly retraining. Prevailing unlearning methods typically rely on suppression-based approaches that lower the probability of specific responses. However, our diagnostic analysis across five baselines reveals that much of the suppressed probability mass consistently flows toward the model's original top- preferences, which may continue to reveal forgotten knowledge through similar expressions. We identify that the root cause is the normalization in softmax, whereby suppressed probability must redistribute but current objectives provide no mechanism to direct this flow. To address this, we propose Safe-Anchored Relocation (SAR), which moves beyond suppression by explicitly relocating probability mass at the token level. SAR constructs dynamic reference branch and current model token sets at each answer position, then employs an InfoNCE-like loss to relocate probability toward reference defined anchors. Both theoretical analysis and experiments on diverse benchmarks confirm the effectiveness of our approach.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.