The Shape of Forgetting: Modeling Retrained Outputs for Machine Unlearning
Abstract
Machine unlearning is commonly evaluated against an oracle model retrained without the data to be forgotten. In this paper, we empirically investigate whether the outputs of such a retrained model on forget samples can be approximated using only the original model and forget data. This is particularly relevant when retain data, i.e., the original data excluding the forget samples, are no longer available. We study recurring patterns in the transformation from original to retrained outputs by decomposing it into four questions: which forget samples change argmax, which class loses probability mass, how much mass is removed, and where it is redistributed. We observe that uncertainty predicts argmax-change propensity, removed mass concentrates on the original top-1 class in high-accuracy regimes, gained mass favors high-ranked alternatives, and even argmax-stable samples require calibration. Based on these observations, we construct synthetic retrained targets and distill an unlearned model directly on forget data. An extensive set of experiments on Food-101, Oxford Pets, Stanford Cars, CIFAR-100, Tiny-ImageNet with ResNet-18, DeiT-T, and ViT-S evaluate the approach in forget-only regimes with different fractions of removed samples. We show that an unlearned model distilled from the proposed synthetic targets achieves a favorable tradeoff between forgetting, utility, privacy, and distributional alignment: across random 10% and 20% forgetting settings, it tracks the retrained model more closely than competing forget-only baselines in KL/TV on forget samples, remains competitive in test accuracy and membership-inference evaluation, and is substantially faster than retraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.