acceptodds
Under review as a conference paper at ICLR 2027

Neighbourhood-Robust LLM Unlearning in Query and Weight Space

Abstract

Unlearning is a proposed safeguard against hazardous or private knowledge, but it protects only if adversaries cannot recover what was removed. Objectives such as negative preference optimisation (NPO) suppress the answer only under the forget questions and at the returned weights, whereas attackers rephrase and follow up, and deployed models are fine-tuned. We cast unlearning as adversarial robustness over the queries reaching a fact and the weights further fine-tuning reaches, and propose neighbourhood-robust unlearning (NRU): a red-teaming probe generator, refreshed as the model changes, reads its responses to find queries that still extract each fact, and each update lowers a soft maximum of the loss over these queries, optionally after a simulated fine-tune. On a relation-level split of TOFU, ten probe strategies with a detector checked against a retraining oracle recover 0.83 of the answerable facts from the original model and 0.08 from the oracle, and NPO still leaves 0.51 recoverable at suppression 0.37 on the direct question. At matched utility, with additional training, a query neighbourhood lowers recovery by up to 0.345 (0.260 for NRU's generator arm), against 0.085 for mining leaked responses; NRU's refresh and soft maximum each lower recovery at 32 epochs. Simulating fine-tuning exposes a trade-off between the axes. Likelihood checks at the forget prompt thus give false assurance; safety audits should measure what an adaptive adversary can still extract, which training over query neighbourhoods reduces.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.