acceptodds
Under review as a conference paper at ICLR 2027

Why Erased Embeddings Look Safe: Fitting the Eraser on the Rows You Release

Abstract

Concept erasure is usually evaluated by how far an attacker's accuracy on the erased representation sits above chance. We show that this number depends on two choices that papers rarely report: which family of attackers runs the probe, and whether the eraser was fitted on the same rows that are released. Owners typically fit the eraser on the data they hold, which is also the data they release. We prove that when the embedding has at least as many dimensions as released rows, this fit subtracts each row's own class centroid, and that for any number of rows it leaves a probe trained on one part of the release pointing against the rest. A sweep over seven embeddings, six ratios of rows to dimensions and three erasure methods confirms both consequences. Linear probes fall below chance at all 28 settings, which the standard metric reads as strong erasure. On ReLU embeddings, where more than half of the coordinates are exactly zero, the same fit turns every zero into a class-specific constant, and a depth-4 tree recovers the protected attribute at 0.79 above chance while linear probes sit 0.20 below it. Fitting the eraser on disjoint rows removes both effects. The mechanism also explains why, on 109 evaluation cells with three published erasure methods at matched utility, two attacker families of equal size disagree about whether to release on 9.1% of cases. The disagreement recurs on 73 new cells, and it extends to method rankings: a kernel-based eraser leaks less than LEACE under one attacker family and no less under the other.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.