acceptodds
Under review as a conference paper at ICLR 2027

Explicit Representation Alignment via Subspace Elimination for LLM Unlearning

Abstract

Large Language Models (LLMs) may retain targeted knowledge, including sensitive or hazardous information, even when such knowledge is undesirable in deployment. Machine unlearning aims to suppress the influence of selected data while preserving useful model behavior. We present ERASER, a representation-guided unlearning method that combines a data-derived subspace target with a ranking-based Disentangle-Head. ERASER constructs a fixed aggregate forgetting projector offline from predefined forget groups and uses it to form a sample-specific residual target during training, while ranking encourages higher learned scores for retain samples than for forget samples. We evaluate ERASER on WMDP, TOFU, MUSE, and Fictional Knowledge, observing favorable forgetting-retention trade-offs relative to the evaluated baselines. On Fictional Knowledge, it reduces recovery on exact, paraphrased, and compositional probes. Under the tested protocols, ERASER also yields lower jailbreak attack success rates than the pre-unlearning model and RMU. We additionally evaluate forget-set membership inference and sensitivity to limited-data relearning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.