Does Removing a Prediction Remove Its Encoding in Language Models?
Abstract
To erase harmful or protected content, unlearning often suppresses a prediction and assumes the encoding behind it disappears together. A surviving encoding keeps that content recoverable. Yet it is not well understood whether this encoding, the internal distinction between tokens, needs its own prediction to form or persist. We measure it as identity dispersion, the spread among mean representations before each token of a category, such as a part of speech. We test this causally in four models up to 1.5B parameters, removing the prediction in four ways, from scratch or after pretraining, against identical unconstrained twins. Removal provably raises prediction loss by at least what the prediction gained from context, so the prediction is gone. In every removal, at least half the categories keep 94% or more of the twin's dispersion one layer below the output, and with one exception no category's median falls below half. Trained without the prediction, models stay 58 to 73% as decodable by a linear probe as their twins. Across categories, the fraction of variance separating a category's tokens tracks the information its prediction uses, and a law we fit on four models predicts it in six other model families. Removing a prediction does not necessarily remove its encoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.