Certified Retain Degradation Control for Counterfactual Guided Machine Unlearning
Abstract
Machine unlearning can introduce new prediction errors on retained data, and searching over update configurations can amplify this risk through selection bias. Reliable deployment therefore requires a degradation guarantee that remains valid after configuration selection. We propose conformal machine unlearning configuration evaluation (CMUCE), a unified framework for certification and selection that combines deletion-efficacy screening with finite-sample control of positive prediction regret relative to a fixed reference model. For each realized deletion request, we establish a guarantee that every accepted configuration has retained-population risk below a prescribed tolerance, even when candidates share calibration data and the final configuration is chosen from the accepted set. The construction combines Hoeffding–Bentkus bounds with family-wise testing and accommodates any calibration-independent candidate generator. We further characterize when counterfactual guidance improves the certificate: under explicit anchor-quality, gradient-alignment, and local score-transfer conditions, a paired one-step comparison yields a certificate-ordering bound whose calibration and anchor-yield error terms decay exponentially with their respective sample budgets. Under a Hellinger shift budget , an additional risk margin of extends certification to the deployment distribution. Experiments on vision, text, and tabular benchmarks show that guided candidates can attain lower retain regret and tighter certificates, while joint selection identifies configurations meeting both deletion and retain-risk requirements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.