Nonmember-Law Matching for Suppressing Unlearning Signatures
Abstract
Machine unlearning can suppress a language model's knowledge of forget data while leaving it vulnerable to membership inference: forget examples remain distinguishable from data the model was never trained on, revealing which records were used in training or targeted for removal. Existing LLM unlearning objectives lower forget likelihoods or redirect internal representations without reference to nonmember behavior, so unlearning can stop short or overshoot, and both leave identifiable signatures. We propose Nonmember-Law Matching (NML), which uses the original model's likelihoods on nonmembers as the target for forget examples. NML matches the distribution of forget-example negative log-likelihoods to this nonmember reference and preserves retained behavior through distillation, without requiring a retrained model. On two unlearning benchmarks, TOFU and MUSE-News, NML substantially reduces unlearning signatures under multiple likelihood-based attacks while maintaining forgetting and utility comparable to or better than strong baselines, and question-only identification attacks provide further evidence of this reduction. These results suggest that calibrating against nonmember behavior is a simple and effective way to reduce the identifiable traces of unlearning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.