acceptodds
Under review as a conference paper at ICLR 2027

ATR: Shadow-Free, Retrain-Free Evaluation of Machine Unlearning

Abstract

Machine-unlearning metrics such as TOFU Forget Quality and privacy-leakage scores do not measure forgetting in absolute terms. Each compares a signal on the forget set, such as its loss or truth ratio, to a reference (a retrained model, the retain set, or a held-out population) and reads the result against a fixed threshold. We show this reading is unreliable. On TOFU, running one method at five random seeds, with method, epoch, and data held fixed, yields Forget Quality values spanning three orders of magnitude, so a threshold verdict, and any ranking built on it, largely reflects the seed. We trace this instability to the metric's -value read-out. We then propose ATR (Artifact-Trace Reduction), an evaluator that trains no shadow models and no retrained oracle. ATR reuses the unlearning run one already performs: it reduces each example's unlearning trajectory to a compact descriptor and scores a method by a two-sample distance between the forget set's descriptors and an explicitly chosen reference. On CIFAR-10 image classifiers and a 1.5B language model on TOFU, ATR exactly reproduces the method ranking of a per-example shadow-model membership attack, the trusted but expensive gold standard (Kendall ), at roughly 7% of its cost. At 7B, ATR's ranking never flips across seeds, whereas Forget Quality's flips on a third of seed-pairs, and ATR's cross-seed spread is three to nine times tighter. Two limits apply. The score's orientation depends on the regime and must be fixed against a gold standard or calibration in new settings. And the membership-attack validation is at 1.5B, so transfer to 7B is argued rather than measured.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.