acceptodds
Under review as a conference paper at ICLR 2027

Towards Substantive LLM Unlearning: Robust Evaluation and Expression-Insensitive Metric

Abstract

Large language model (LLM) unlearning aims to remove private, copyrighted, or harmful knowledge from trained models while preserving general utility. Despite recent algorithmic progress, reliably determining whether target knowledge has been removed remains challenging. Existing metrics primarily assess immediate post-unlearning behavior, often overlooking robustness and thus becoming susceptible to superficial behavioral changes and unreliable residual-knowledge measurements. This work proposes RUMOR, a comprehensive framework that evaluates unlearning-metric reliability from data-side and model-side robustness perspectives. RUMOR tests whether metrics can reliably distinguish different degrees of residual knowledge across expression forms and remain stable under model-level red-team unlearning attacks. We further introduce Unlearning Depth (UD), a robustness-oriented metric that measures whether target knowledge is consistently forgotten across expression forms. Across diverse unlearning tasks and methods, existing metrics can be misled by superficial behavioral changes, yielding unreliable residual-knowledge measurements. These findings reveal critical limitations in current unlearning evaluation and provide an extensible framework for assessing whether unlearning genuinely removes target knowledge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.