Caffeine: Machine Unlearning via Curvature-Aware Gradient Ascent
Abstract
Deleting data from a trained model without retraining requires an update that cancels the deleted data's contribution. Influence functions specify this update via the inverse Hessian, but evaluating it requires iterative passes over the training data, which is often no longer available when a deletion request arrives. Existing approximations read the retained data, restrict the change to the final layer, or drop curvature rescaling altogether. We present Caffeine, a source-free unlearning algorithm that operates as a curvature-aware gradient ascent. Since the update moves along the gradient, it needs the curvature strictly along that one direction—a single scalar that two forget-set gradients evaluated a small step apart determine without any exact second-order computation or retained data. Caffeine measures this scalar at each saved training checkpoint and aggregates the resulting curvature-scaled steps into a single update to every layer of the deployed model. Across class and random-subset deletion on five image and text datasets, Caffeine is the only approximate method that both forgets and preserves accuracy without retained data, running up to faster than retraining on image tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.