acceptodds
Under review as a conference paper at ICLR 2027

HOW ERASED IS ERASED? QUANTIFYING AND IMPROVING THE RECOVERABILITY OF CONCEPT REMOVAL IN CLIP

Abstract

Concept-erasure (“unlearning”) methods for vision-language models such as CLIP typically report success by showing that zero-shot accuracy on the erased class collapses to near zero. We argue this metric conflates two different things: information being removed from the representation versus information being merely hidden from one particular classifier (the zero-shot text-alignment rule). We introduce a cheap evaluation protocol – training a k-shot (k=5, 20, 50) linear probe directly on the post-erasure embeddings and measuring how much of the erased class it recovers – and apply it to three erasure mechanisms on frozen CLIP ViT-B/32 features: a naive single-direction baseline, an iterative multi-direction method (INLP), and a data-free, text-embedding-derived nullspace-projection method whose mechanism matches a recently published approach (Mishra et al., 2025). Across three image domains (CIFAR-100, EuroSAT, and Food101), we find that (i) every method we test leaves the erased concept substantially recoverable despite 0% zero-shot accuracy, and (ii) the data-free, text-derived method is consistently the most recoverable of the three (92–97% recall at k=50), more so than the naive image-based baseline, on all three datasets. On CIFAR-100, where class hierarchy labels are available, we show the leaked signal is not unstructured noise: attack false positives concentrate in the same coarse superclass as the erased class 9.8× more than chance for the text-derived method (measured over 100 randomly-sampled target classes), suggesting it erases a shared semanticdomain direction rather than a class-specific one. We also find that whether this method’s extra recoverability comes with an extra utility cost is entangled with how strong the zero-shot baseline is to begin with: the two datasets with a wellseparated zero-shot baseline (CIFAR-100, 61.7%; Food101, 84.3%) both show the text-derived method costing the most accuracy on retained classes, while the one dataset with a weak baseline (EuroSAT, 40.1%) shows the opposite – which we argue is an artifact of that weak baseline destabilizing the metric, not a genuine reversal of the phenomenon. We propose Ensemble Subspace Projection (ESP), a bootstrap-aggregated multi-direction erasure operator, which reduces attack recall relative to the classical iterative baseline at equal removed-dimension budget on every dataset we test (by 1–9 recall points), for free on CIFAR-100 but at a small utility cost on the other two; no method we test achieves what we would consider genuine erasure. All of the above uses CLIP ViT-B/32; we additionally confirm our two central findings – NSP-text is the most recoverable and the most utility-costly method – hold essentially unchanged on the 2.8× larger ViT-L/14 backbone. Finally, extending from object classes to a demographic attribute on FairFace, we find naive erasure of a binary gender attribute leaves recoverability that swings 8.9× across race groups (7.7–68.6% recall at k=50) even though every group is erased and attacked identically – a class-level mechanism producing the same kind of group-uneven outcome documented for retraining-based forgetting, and a caution that fairness properties of an erasure method cannot be read off its object-classification behavior alone; tellingly, ESP – our own defense – leaves this disparity just as wide (if not wider) even where it reduces average attack recall elsewhere. We discuss the weak-baseline pitfall and other open problems for future work on evaluating unlearning in vision-language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.