Batch-Norm Statistics: The Unrecorded Convention in Unlearning Audits
Abstract
Unlearning benchmarks release checkpoints with tables of accuracies and pass/fail verdicts, and a reader treats each entry as a property of the released checkpoint. A batch-normalized checkpoint, however, also ships normalization statistics fitted from data. An audit's retrained reference fits them on retained data by construction, while the unlearned checkpoint ships whatever its own run left, and no release we audit records how. We show that the published numbers of such checkpoints, and some verdicts, also depend on this fitting convention. We reproduce each checkpoint's published numbers and then refit only its statistics on retained data at byte-identical weights. The refit raises forget accuracy beyond the spread MU-Bench reports across training seeds on 47 of the 221 reproducible MU-Bench checkpoints we test. A method's average does not predict which ones. In a targeted test, removed data in the statistics does not explain which checkpoints move. Releases should therefore state how their statistics were fitted, and a reader can check one checkpoint on a single GPU in under an hour.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.