U2E: Protected Common Descent for Targeted Behavioral Factual Unlearning
Abstract
Targeted factual unlearning seeks to suppress a small set of obsolete or sensitive associations in a deployed language model while preserving its remaining knowledge and capabilities. Knowledge editing offers a local intervention, but its formulation around a replacement value does not determine whether one protected update can suppress the old fact across several prompts and internal representations whose required directions may conflict. Building on the key–value view of transformer memory and covariance and null-space protection, we formulate factual unlearning as protected common descent. This formulation characterizes whether multiple access-specific decreases admit a shared protected update and, when feasible, the minimum cost of satisfying the local constraints. The associated quadratic problem has a unique solution with an active-support rank bound. Based on this formulation, U2E constructs low-rank candidate updates using local progress estimates and validates them on the full nonlinear model before commitment. Experiments demonstrate favorable tradeoffs between old-answer suppression and measured retention under the evaluated protocols, supporting protected common descent as a practical framework for selective factual unlearning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.