What Does It Mean to Unlearn? Indistinguishability and Admissibility
Abstract
LLM unlearning is increasingly required to remove sensitive information without retraining large models from scratch, and is often considered successful once the model stops revealing that information. However, a model that no longer outputs this information may still reveal it through its preferences among other answers, while refusal or degenerate outputs give up useful behavior. A few studies detect such failures, and none defines or achieves success. We argue that a truly unlearned model should leave no trace of having known the target information, while still acting naturally without it. We therefore require indistinguishability of outputs across worlds that differ only in the target information, leaving no evidence of the target. To keep responses appropriate, we further require admissibility, specified by an admissible region of response distributions shared across worlds. Based on this region, we introduce WorldLeak and prove that bounding it certifies -indistinguishability at a chosen resolution without a retrained reference model. For repair, the closest distribution within a WorldLeak budget gives a unique target. WorldRepair realizes it by reweighting response likelihoods and adapting erasure strength to residual leakage under a retain constraint. On TOFU, MUSE, and WMDP, WorldRepair matches strong baselines on native metrics with the lowest WorldLeak and cross-world distinguishability. Our framework defines successful unlearning in a form that can be both certified and optimized.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.