acceptodds
Under review as a conference paper at ICLR 2027

Probe-Geometry Alignment: Auditing and Suppressing Representational Memorization in Language Models

Abstract

Unlearning methods for language models are judged by what the model says. Prompt the model with the opening words of a target passage, and if it no longer completes the passage, the content is reported as removed. A model can pass that test and still encode the passage. We measure what is left with a probe that has to generalize. A linear classifier is fitted on the hidden states of several memorized sequences, then tested on a memorized sequence it has never seen. To succeed, it must read a property of memorization itself rather than a particular string. It succeeds on every model we tested. Across three pretrained architectures, the memorization-specific gap runs from +0.19 to +0.35. It is stable across the Pythia family from 70M to 1B parameters, and it holds on a pool whose membership is verified by corpus lookup rather than by likelihood. The direction the probe reads is causally separable from the machinery that recites the text. No behavioral erasure we applied removes it, from five published unlearning objectives to weight editing to imitating a model that never memorized, however far the output moves. That separation is what makes an efficient method possible. Probe-geometry alignment (PGA) constrains one scalar per depth, the projection onto the direction the probe reads, and leaves the rest of the residual stream alone. On Pythia-70M it buys 6.6 times as much recall suppression per nat of capability as the best of four isotropic settings on the same recipe. It also suppresses more recall than the behavioral baseline that retains the most capability, while retaining more capability than that baseline, and it carries over to Mistral-7B. Constraining a readout can also reverse it rather than empty it. Raw probe accuracy then falls below chance while the classes stay separable, so an audit that quotes that number records erasure where none occurred. A multi-position variant drives a fresh mean-pooled linear probe to chance at mid depths, at a further cost in capability. The audit and the method share one direction, so a deployment able to fit a probe can state what its deletion claim is worth, and act on the answer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.