What Survives Unlearning? Calibrated Subject-Level Auditing of Residual Knowledge in LLMs
Abstract
Unlearning for large language models (LLMs) is evaluated almost exclusively through forget-set averages, which cannot tell whether each forgotten subject has become indistinguishable from a model that never saw it. We introduce ReCAUL, a post-hoc, method-agnostic audit that answers this subject-level question at an unchanged query budget. ReCAUL compares an unlearned checkpoint with deletion references trained without the forgotten subjects, calibrates each relation and wording against a reference-only null, subtracts a matched retained subject, and cross-fits two wordings so that one selects each subject's strongest relation and the other confirms it; its randomization test repeats the selection in every draw. On TOFU with Llama-3.1-8B, the resulting statistic, CF-PMAX, exposes 2.6 times the residual found by uniform aggregation of the same ten queries (6.52 vs. 2.54; ) in a task-vector checkpoint that passes the aggregate screen, with positive residual in all 24 audited subject blocks. The result replicates for a gradient-difference checkpoint, on the same relations (6.24; 24/24), holds for negative preference optimization (NPO; 4.79; 22/24), and stays positive under recalibration, re-pairing, and single references, while reference-only pseudo-candidates trigger it at a near-nominal 4.9% rate. Prompted with the audited relation, the NPO checkpoint regenerates more of the forgotten answer than its deletion references, while its retained subjects show no such gap. Code and frozen score rows that reproduce the primary audit are provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.