acceptodds
Under review as a conference paper at ICLR 2027

AGTTU: Attention-Guided Target Token Unlearning for Large Language Models

Abstract

Large language models exhibit strong language understanding and reasoning capabilities, but also memorize private or copyrighted content, raising privacy and compliance concerns. Existing unlearning methods often uniformly suppress entire forget samples, inadvertently degrading general language capabilities unrelated to the unlearning targets and causing excessive forgetting. Token-level approaches alleviate this issue, but some rely on auxiliary models or additional perturbation-based scoring. We propose Attention-Guided Target Token Unlearning(AGTTU), which selects target tokens based on the attention they receive from subsequent answer positions, using the model's own attention allocation to guide targeted unlearning. We further introduce Contrastive Logits Preference Loss (CLPL) to suppress target tokens while limiting damage to non-target content, together with a retain-set objective to preserve model utility. Experiments on TOFU and MUSE Books demonstrate that AGTTU improves the trade-off between unlearning effectiveness and model utility. On TOFU with Llama-3.2-1B, AGTTU achieves higher reported forget quality and model utility than the listed baselines across all three forgetting ratios. On MUSE Books, compared with TPO, AGTTU reduces the verbatim memorization score on the forget set by 4.51% while increasing the knowledge memorization score on the retain set by 15.17%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.