Conformal Unlearning: A New Paradigm for Output Control Under Conformal Prediction
Abstract
Existing machine unlearning definitions generally constrain the objective of unlearning to approximating a model retrained from scratch after removing the forget set. However, this formulation equates unlearning to `untraining' under all scenarios, presenting substantial practical limitations as it often neglects the gap between parameter-space proximity and output-level consistency, is vulnerable to falsification, and overstates retrained models as a universal gold standard for unlearning. To address these issues, we introduce conformal machine unlearning, a new paradigm that does not require a retrained model as the reference, and instead focuses on controlling the output of the model. We formalize conformal unlearning as an output-control framework that enforces high conformal coverage on retained data and high conformal miscoverage on forgotten data, develop empirical evaluation metrics, and present an algorithm that optimizes these conformal objectives. Extensive experiments on vision and text benchmarks verify that the proposed method effectively removes targeted information while preserving utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.