acceptodds
Under review as a conference paper at ICLR 2027

CROP: Curvature-Restricted Orthogonal Projection for Retrain-Faithful LLM Unlearning

Abstract

Machine unlearning removes the influence of specific training data from a language model without retraining. Existing methods choose their update direction without reference to the joint forget and retain curvature and are graded by suppression, which rewards overshooting the retrained model. We present CROP (Curvature-Restricted Orthogonal Projection), built on the measured premise that memorization concentrates in parameter directions sharply curved under the forget loss and flat under the retain loss. CROP restricts forget ascent to this curvature-selected subspace, projects the step orthogonal to the retain gradient, anchors on retain descent, and selects its operating point by the distance between its forget profile and the retrained model's, or by a calibrated oracle-free rule. On TOFU, under a leakage-free protocol that selects checkpoints on validation authors and scores held-out ones, CROP leads RMU, the strongest baseline, on held-out forget quality on two of three seeds at a utility cost; as an analysis upper bound, its oracle-best checkpoint lies within KS – of the retrained model on every seed, close to the – spread between independent retrains. On MUSE it lands closest to the retrained reference for two 7B families, and on WMDP, which has no retrain oracle, a supplementary suppression test shows it suppresses hazardous knowledge further than RMU while remaining functional. Under full fine-tuning, LoRA, and partial-forget relearning attacks it recovers no faster than the attacked retrain oracle, while RMU, NPO, PDU, and WGA rebound to near pre-unlearning knowledge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.