Erase, Don’t Distort: Machine Unlearning through Representation Realignment
Abstract
Machine unlearning is often carried out by deliberately increasing the loss on the forget set to suppress the model’s response to targeted samples. While effective at inducing forgetting, this can leave a distinct statistical footprint that is vulnerable to membership inference attacks, a phenomenon known as the Streisand effect. The core issue is that loss maximization specifies what should be erased without controlling how the underlying representations evolve, often driving forgotten samples into abnormally high-loss regions that are easy to distinguish from naturally unseen data. We instead formulate unlearning as a controlled subspace realignment process that selectively modifies forget-specific dimensions while preserving the latent structure shared by retained data. To realize this formulation, we first isolate a low-dimensional subspace that captures the dominant discrepancy between forget and retain representations, while constraining its orthogonal complement to remain close to the original model. Within this subspace, we then realign the forget distribution toward a nearby retain distribution by matching both its centroid and its dispersion, thereby avoiding the abnormal concentration induced by destructive forgetting. This design confines forgetting to the representation directions where it is needed while preserving shared structure elsewhere. Comprehensive evaluations across CIFAR-10, CIFAR-100, SVHN, Tiny ImageNet, and ImageNet32-200 demonstrate a strong privacy–utility trade-off, driving membership inference performance close to the random-guess regime () while matching or surpassing representative baselines in retain utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.