Unreal Unlearning: Behavioral Forgetting Without Knowledge Erasure
Abstract
Machine unlearning aims to remove targeted knowledge from LLMs, but current benchmarks mainly test whether the model still gives the correct answer, not whether the underlying knowledge is gone. We study unreal unlearning, where behavioral forgetting occurs while task-relevant information remains linearly decodable from hidden states. Across eight unlearning methods on the WMDP-Bio domain, accuracy falls to near chance while substantial internal evidence for the correct answer remains. We introduce Knowledge Lens, a framework which evaluates sets of True/False claims, and reads two signals, output-logit margin and a linear probe on the hidden state. Instantiated with WMDP-Bio, we find that % of questions are suppressed, with external evidence at or below chance while internal evidence remains above chance, for instance the gap between WMDP benchmark and the averaged internal measure is 60.8 percentage points. We vary unlearning methods and their checkpoints, three additional model architectures, an independently released nine-model method suite, all using models released by the original works, and extend to the WMDP-Cyber domain. In all these settings, Knowledge Lens consistently reveals substantial internal-external separation even when behavioral performance approaches chance. These results show that behavioral forgetting can substantially overstate knowledge removal and that unlearning should be evaluated through internal, interpretable approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.