Label-Level Certified Untraining for Privacy Guarantees in Deep Neural Networks
Abstract
Untraining removes the influence of specific training data (the forget set) from a model without retraining from scratch, a practical necessity for privacy. Yet no method guarantees privacy for deep neural networks (DNNs), where untraining matters most due to excessive costs: heuristic methods offer only empirical evidence, and certified methods, which draw on differential privacy (DP), yield vacuous guarantees as model size grows. We introduce -certification, the first label-DP-style formalization of untraining, which certifies the model's confidence scores on the forget set rather than its parameters, keeping guarantees meaningful at scale. We then propose certified untraining via prediction calibration (CUP), which calibrates these scores toward a retraining-free reference, the original model's confidences on unseen data. CUP transfers across settings, even without access to the remaining training data, by leveraging this retraining-free reference. Extensive ablations show that CUP can yield privacy guarantees without compromising untraining performance and model generalization. On ViT-Large, a scale unexplored by certified work, CUP guarantees that any test distinguishing untrained from retrained forget-set confidences errs at least of the time (random guessing: ).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.