Auditing Fact Deletion in Language Models with Separable Memory
Abstract
Language models with separable memory aim to decouple source-specific knowledge from general world knowledge, language understanding and reasoning capabilities. Unlearning then becomes relatively straightforward: simply delete the memory units associated with the targeted knowledge. We show that a naive comparison of model outputs is an unreliable measure of unlearnability, and propose a methodological audit for such systems that measures the effectiveness of memory removal. We apply the audit to Continuous-Query Limited Memory Language Models (Co-LMLM) and Natively Unlearnable Large Language Models (NULLs) on five factual-recall benchmarks. Much of the post-deletion correctness in both systems is also present when separable memory is disabled. This is expected for NULLs by design, but less so for Co-LMLM, which aims to keep facts out of its weights. In NULLs, the backbone recites up to % of audited facts without any source memory, and removing more sources does not remove more facts but damages up to % of nearby ones. For Co-LMLM, up to % of all facts that pass a naive accuracy test as unlearned are still recited by the backbone. A single entry re-inserted into Co-LMLM's memory restores -% of facts the model no longer answered after deletion. We propose the audit as a protocol that future work on separable memory can use to support deletion claims.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.