acceptodds
Under review as a conference paper at ICLR 2027

U2E: Protected Common Descent for Targeted Behavioral Factual Unlearning

Abstract

Targeted factual unlearning seeks to suppress a small set of obsolete or sensitive associations in a deployed language model while preserving its remaining knowledge and capabilities. Knowledge editing offers a local intervention, but its formulation around a replacement value does not determine whether one protected update can suppress the old fact across several prompts and internal representations whose required directions may conflict. Building on the key–value view of transformer memory and covariance and null-space protection, we formulate factual unlearning as protected common descent. This formulation characterizes whether multiple access-specific decreases admit a shared protected update and, when feasible, the minimum cost of satisfying the local constraints. The associated quadratic problem has a unique solution with an active-support rank bound. Based on this formulation, U2E constructs low-rank candidate updates using local progress estimates and validates them on the full nonlinear model before commitment. Experiments demonstrate favorable tradeoffs between old-answer suppression and measured retention under the evaluated protocols, supporting protected common descent as a practical framework for selective factual unlearning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.