acceptodds
Under review as a conference paper at ICLR 2027

Steepness Is Not Selectivity: Why Post-hoc Unlearning Misses the Retraining Direction

Abstract

Machine unlearning is usually framed as a search for a better objective. That framing rests on a premise that is rarely tested: that the update retraining would make is one a post-hoc procedure can find. We test it directly. Using as an oracle a model retrained without the forget set from the same base, we find the obstacle is not entanglement: the straight path to that oracle is almost perfectly selective — on our primary family forget loss rises 140× while retain loss rises by at most 1.19× and ends below where it started — yet standard unlearning updates that reach a comparable forget level are orthogonal to it, at cos ≈ 0.009 in a space where two runs of the same method aimed at disjoint fact sets align at 0.035. The mechanism is a separation between steepness, forget loss gained per unit of weight movement, and selectivity, forget loss gained per unit of retain damage: gradient descent maximizes the first, unlearning wants the second, and nothing ties them together. Plain SGD, where the steepest-descent argument applies literally, is steeper still and barely selective at all, so the preconditioning in the optimizer people actually use is a partial remedy rather than a confound. Preconditioning ascent by the retain curvature instead, as the mechanism prescribes, cuts collateral damage by an order of magnitude and moves further from the oracle direction: the local optimum of selectivity is not the retraining direction. Two further results say the premise fails even if the search succeeded. Being better aligned does not help: weight arithmetic is 7–8× closer to the oracle direction and no more selective, and across the four methods alignment does not track selectivity. And the target does not transfer: measured from a different full-data solution the retraining displacement is nearly orthogonal (cos = 0.029) to the one measured here, so there is no single vector to compute once and apply. The existence claim holds in three model families spanning 0.5–1.5B, and the geometry in the two where we ran the full comparison.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.