acceptodds
Under review as a conference paper at ICLR 2027

Certified Unlearning for Distilled Teacher-Student Models with a Black-Box Teacher

Abstract

We study certified machine unlearning for student models distilled from a teacher accessible only through black-box prediction queries. After deployment, the student must unlearn requested samples, while the teacher is unavailable for retraining and accessible only through black-box prediction queries. We introduce a student-side certification framework that re-distills the student on the retained data using only black-box teacher access, and releases a randomized model by carefully calibrating the perturbation to certify indistinguishability against the ideal but counterfactual retrain-from-scratch reference, where the latter would require retraining the (inaccessible) teacher without the forget samples. Our certificate builds on a student sensitivity bound for distillation, which we obtain by characterizing how the teacher's predictions would change under retraining without the forget samples, and then propagating this bound through the distillation objective to bound the student's parameter shift. The final student sensitivity is used to perturb the released student parameters with calibrated Gaussian noise to ensure the target certification guarantees. In doing so, our work provides a first study of unlearning certification for black-box teacher-student distillation. While our work is primarily theoretical and relies on several regularity assumptions for tractability, we hope our results will provide interesting insights for unlearning in this increasingly common regime.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.