acceptodds
Under review as a conference paper at ICLR 2027

Forget, But Not Too Much: Self-Limiting LLM Unlearning

Abstract

Machine unlearning aims to remove selected knowledge from a trained model without retraining from scratch. Most existing LLM unlearning methods reduce the probability of the target answers. This does not match how TOFU evaluates forgetting. TOFU compares the unlearned model with a model that never saw the forget set, so pushing the target answers far below that reference can reduce Forget Quality just as retaining them can. We show that common one-sided objectives have no stationary point at this target. They can therefore keep updating after the model has reached its best agreement with the reference. We introduce NGREF, a self-limiting objective that matches the model to its frozen pretrained checkpoint on the forget set while retaining performance on the remaining data. The pretrained checkpoint becomes a fixed point of the forget objective. On the full published TOFU comparison grid with two models, three forget splits and four LoRA ranks, NGREF exceeds the best Forget Quality reported by nine published methods in all 24 settings. A controlled objective swap raises Forget Quality from approximately zero to 0.585 while leaving model utility similar. In a training sweep, a strong one-sided baseline reaches 0.766 and then falls to 0.0013 even though its forget loss continues to improve. Reference matching instead reaches 0.99 and remains close to that value over the measured range. These results show that the endpoint of unlearning matters as much as the strength of forgetting. Code: https://anonymous.4open.science/r/ngref-unlearn-4F8E/README.md

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.