Explicit Representation Alignment via Subspace Elimination for LLM Unlearning
Abstract
Large Language Models (LLMs) may retain targeted knowledge, including sensitive or hazardous information, even when such knowledge is undesirable in deployment. Machine unlearning aims to suppress the influence of selected data while preserving useful model behavior. We present ERASER, a representation-guided unlearning method that combines a data-derived subspace target with a ranking-based Disentangle-Head. ERASER constructs a fixed aggregate forgetting projector offline from predefined forget groups and uses it to form a sample-specific residual target during training, while ranking encourages higher learned scores for retain samples than for forget samples. We evaluate ERASER on WMDP, TOFU, MUSE, and Fictional Knowledge, observing favorable forgetting-retention trade-offs relative to the evaluated baselines. On Fictional Knowledge, it reduces recovery on exact, paraphrased, and compositional probes. Under the tested protocols, ERASER also yields lower jailbreak attack success rates than the pre-unlearning model and RMU. We additionally evaluate forget-set membership inference and sensitivity to limited-data relearning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.