acceptodds
Under review as a conference paper at ICLR 2027

Predicting the Collateral Cost of Linear Concept Erasure

Abstract

Linear concept erasure removes a target attribute from a frozen representation and guarantees linear guardedness: the optimal affine least-squares predictor of that attribute can no longer improve on the constant predictor. The guarantee says nothing about what the removal costs other attributes, and that cost is normally established only after the fact, by retraining a probe for every retained attribute. We show that the corresponding loss in linear accessibility can instead be computed in closed form beforehand. For any linear map that guards the concept set, the post-erasure accessibility matrix of the retained concepts is bounded above by a Schur complement of the accessible task covariance , whose diagonal entries are the linear of each concept and whose off-diagonal entries are the linearly accessible coupling between concept pairs, with equality for minimal-kernel erasers including LEACE. The Schur term is therefore the minimum collateral loss that guarding makes unavoidable, and for LEACE the leverage profile of the erased subspace additionally locates the edit. Across frozen representations and four datasets, both closed forms hold to within numerical error. Erasing every concept of every dataset in turn, the accessible coupling read off predicts held-out AUROC degradation of retrained probes better than squared label correlation does, and screening all CelebA triples from alone separates the cheapest erasures from the most damaging before any eraser is run.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.