Attention-Only Machine Unlearning in Large Language Models
Abstract
Machine unlearning aims to remove designated knowledge from a model while preserving its utility and minimizing computational cost. However, existing approaches often struggle with *knowledge entanglement*, where targeted and retained knowledge share model representations. In this work, we investigate attention-only unlearning for large language models (LLMs), focusing on a component that is central to the Transformer architecture yet often overlooked as a target for unlearning. Inspired by recent insights into memory retrieval through attention mechanisms, our method systematically identifies context-specific "anchored terms” and confines parameter updates to attention layers to enable targeted, granular unlearning. Experiments show that our approach achieves unlearning effectiveness competitive with existing techniques while preserving performance on retained data and reducing computational requirements. These efficiency gains stem from updating attention modules, which contain fewer parameters than the corresponding multilayer perceptron (MLP) modules in the evaluated architectures. Our findings highlight attention as a promising target for efficient unlearning and offer insights into its role in selectively accessing and suppressing knowledge in LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.