Samplewise Retraining Compatibility: Outcome-Level Evaluation and Guidance for Machine Unlearning
Abstract
Approximate machine unlearning is commonly evaluated by fidelity to exact retraining. Existing evaluations either compare a realized unlearned model against a single retraining reference or compare repeated unlearning and retraining outcomes distributionally. The former reduces stochastic retraining to a single reference realization, while the latter characterizes repeated outcomes rather than the single model produced for an individual deletion request. This leaves an outcome-level question: how should a realized unlearned model be evaluated against retraining's natural variability? We formalize samplewise retraining compatibility (SRC) to evaluate whether each deleted sample's realized behavior falls within the range naturally produced by stochastic retraining. To operationalize SRC with finite retraining references, we introduce BLORO, which calibrates unlearning deviation against retraining's own variability. To pursue SRC without retraining references, we propose ReCo, which estimates sample-specific proxy targets from difficulty-conditioned nonmember behavior and adaptively controls forgetting toward them. Experiments show that BLORO remains stable with finite retraining references, ReCo achieves strong SRC while remaining competitive on conventional set-level unlearning metrics, and models obtained with ReCo better preserve their post-unlearning behavior on deleted samples during subsequent training on retained data. Together, these results support SRC as a principle for evaluating and guiding approximate unlearning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.