acceptodds
Under review as a conference paper at ICLR 2027

Overshooting the Oracle: Why Zero Forget Accuracy Is the Wrong Target for Entangled Unlearning

Abstract

Unlearning methods for entangled forget sets, where forgotten and retained data share a label, are typically trained to drive forget-set accuracy to zero. However, a model retrained without the forget set (usually considered the reference in unlearning) still retains substantial accuracy on it, given similar examples are present in the retain set. Consequently, these methods that achieve zero accuracy on the forget set do not approximate retraining; they overshoot it, at a cost to both utility and privacy. We present a separability ceiling showing that the zero-accuracy objective is infeasible precisely where the forget and adjacent retain sets are hardest to tell apart. We further argue that class-based unlearning (where the goal is to unlearn a particular subclass, e.g., the aquarium fish subclass of the superclass fish), the standard benchmark, barely tests entanglement. We then propose Swift, which replaces both endpoints of the usual objective with proxies that need no retrained model: the forget-loss distribution is matched to a held-out calibration set, and the retain-loss distribution is matched to a cheap probe fitted on retain data alone rather than to the contaminated original model. Both terms are squared 2-Wasserstein distances, and the forget term penalizes over-forgetting as much as under-forgetting. We evaluate on four datasets spanning image classification (CIFAR-100), face recognition (VGGFace2), and toxicity and social-bias detection in text (ToxiGen, SBIC). On CIFAR-100, Swift cuts the L1 deviation from the retrained reference from 134.62 to 22.37 and the membership-inference deviation from 0.556 to 0.032 compared with Two-Stage, while running up to 64× faster than retraining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.