acceptodds
Under review as a conference paper at ICLR 2027

Erase It Where You Found It: Matched-Position Erasure Lowers Demographic Dependence at a Measured Cost

Abstract

Linear concept erasure removes a direction from a model’s hidden states. LEACE, the standard eraser, guarantees that no linear probe can recover the concept from the edited states. This paper shows that the token positions at which a direction is removed decide what the erasure removes and what it costs. The demographic direction comes from swapping the demographic term of a bias item and measuring how much the answer depends on it. That direction is then erased only at the term’s tokens, in every layer. At those positions the erasure lowers the measured dependence by 11 to 28 percent on four open models. LEACE fitted to the same swap data, and random subspaces, remove nothing at those positions: a probe guarantee is not a guarantee about the answer. An additive steering vector at the same positions does as well as the erasure. At the last prompt token neither edit removes the dependence, and at every token both harm the model. On new BBQ items the erasure removes the demographic term’s contribution to the answer, whether that contribution is a stereotype or the information the question needs. On items in which the term is the answer, accuracy falls from 0.80 to 0.16 on Gemma-2-2B. The evidence spans four models, three benchmarks, a pre-registered replication on new items, and controls for direction, rank, energy, readout and operator. The site of the edit decides what is removed and what is lost. An erasure result should state where the edit acts and what it costs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.