Forgotten Compared to What? A Calibrated Audit of Class Unlearning
Abstract
Machine unlearning aims to make a trained model behave like one retrained without designated data. Whether it succeeds depends on what we measure. A model may stop predicting a removed class even though its scores still distinguish that class. Yet finding such a signal is not enough: a model trained without the class may also recognize its images through generalization. The relevant question is therefore: forgotten compared to what? We introduce a calibrated audit that measures recovery from an unlearned model relative to models trained without the removed class, at a fixed operating point. Our main test, woken recall, measures recall after calibrating a single logit offset on retained data at a nominal 5% false-positive rate. We use the existing classifier and choose the offset using only retained-class examples. In a CIFAR-10 study of 480 checkpoints from 16 fixed unlearning configurations, models with near-zero forget accuracy often have higher removed-class recall than retain-only references matched on retained accuracy. A frozen-feature probe gives a separate, more supervised view and leads to the same conclusion. Recovery has to be read against a stated reference and operating point; forget accuracy alone cannot provide that comparison.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.