Certificates Expire: Durable Unlearning for Large Language Models
Abstract
Machine unlearning removes designated knowledge from a trained language model, with release-time checks assessing whether the requested knowledge erasure has been achieved. Durable unlearning faces a distinctive obstacle: subsequent adversarial fine-tuning moves the released parameters and can recover knowledge that passed those checks. Identical release-time outputs need not fix recovery cost, while resistance to sampled attacks does not certify every trajectory within a budget. To address these, we first establish a separation between release-time behavior and durability, locating the missing information in the geometry of the landing point. We then design PLATEAU, a framework that converts this geometry into a certified budget of fine-tuning updates. Within this framework, Directional Measure Estimation reads the forget-loss margin and directional gradient norm at release together with a uniform negative-curvature bound over a prescribed neighborhood. Barrier Certification then converts valid bounds into a closed-form safety radius and an integer budget of complete updates, covering every admissible trajectory and evaluation objective in the declared class without running attacks. In addition, Certified Durability Training uses the radius dependencies to regularize existing unlearning objectives, followed by re-evaluation of the updated release. Under a common declaration on TOFU, MUSE and WMDP-bio at matched forgetting and utility, PLATEAU attains the largest certified budget and the longest recovery on TOFU and MUSE, and every one of the 51 compared method–benchmark cells stays unrecovered within its budget. Together, the separation and certificate make durability a property of the released parameters: any unlearning procedure can be assessed through the same geometric quantities, with a guarantee expressed in the complete updates an adversary can take.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.