acceptodds
Under review as a conference paper at ICLR 2027

WADE: Wasserstein Distributional Alignment for Knowledge Erasure in LLMs

Abstract

Machine unlearning for large language models aims to remove the influence of specified forget data while preserving retained knowledge. Retraining on the retained set is the gold standard, but is prohibitively expensive. Existing approximate methods typically optimize proxy forget losses or steer forget activations toward a single target vector, which can suppress forgotten answers without matching the distribution of forget-query behavior. We propose WADE, a closed-form activation-editing method that casts LLM unlearning as distributional alignment in activation space. WADE edits selected MLP down-projection matrices so that the mean and covariance of forget-query activations match a target distribution induced by unknown-entity prompts, under a second-order Wasserstein objective with a retain-key regularizer. A Bures–Procrustes reduction yields a non-iterative, ridge-regularized weight update that provably does not increase the surrogate objective. On TOFU across three LLM families, WADE attains the lowest KS distance to the retrain-only reference among the evaluated baselines at their published settings; under per-method hyperparameter sweeps it provides the low-KS end of the alignment–utility frontier, while NPO attains higher utility and K-FADE reaches a neighboring point at much higher cost. Membership-inference probes show retrain-comparable aggregate leakage, with a remaining gap under Min-K++.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.