AGTTU: Attention-Guided Target Token Unlearning for Large Language Models
Abstract
Large language models exhibit strong language understanding and reasoning capabilities, but also memorize private or copyrighted content, raising privacy and compliance concerns. Existing unlearning methods often uniformly suppress entire forget samples, inadvertently degrading general language capabilities unrelated to the unlearning targets and causing excessive forgetting. Token-level approaches alleviate this issue, but some rely on auxiliary models or additional perturbation-based scoring. We propose Attention-Guided Target Token Unlearning(AGTTU), which selects target tokens based on the attention they receive from subsequent answer positions, using the model's own attention allocation to guide targeted unlearning. We further introduce Contrastive Logits Preference Loss (CLPL) to suppress target tokens while limiting damage to non-target content, together with a retain-set objective to preserve model utility. Experiments on TOFU and MUSE Books demonstrate that AGTTU improves the trade-off between unlearning effectiveness and model utility. On TOFU with Llama-3.2-1B, AGTTU achieves higher reported forget quality and model utility than the listed baselines across all three forgetting ratios. On MUSE Books, compared with TPO, AGTTU reduces the verbatim memorization score on the forget set by 4.51% while increasing the knowledge memorization score on the retain set by 15.17%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.