OBSERVATION AND TRANSPORT LIMITS OF SOURCE-ONLY CONCEPT ERASURE UNDER DOMAIN SHIFT
Abstract
Concept erasure is often evaluated by whether a chosen probe can recover a nuisance variable from the resulting representation. This criterion does not establish whether a detector trained on source domains will retain its operating point on an unseen domain. Across pathology and natural-image benchmarks, we find that source-fitted linear erasure can remove the observed domain signal while leaving target false-positive rates largely unchanged. We explain this separation through two limitations. First, domain means expose only the content-averaged nuisance response; content-dependent responses can cancel under averaging, and finite-sample grouped means can misleadingly appear to recover them. Second, even perfect knowledge of the source nuisance geometry need not identify the target-optimal erasure subspace without assumptions connecting source and target domains. Target-oracle controls distinguish this transport limitation from limitations of the projector class. Finally, we compare adaptation strategies under two deployment objectives: false-positive-rate calibration and discrimination at matched false-positive rate. Direct threshold adaptation is effective for the first in the studied settings, although small adaptation sets have substantial uncertainty; feature-space adaptation can improve the second by changing the ordering of examples. These results support choosing an adaptation target according to the deployment objective rather than treating source-side erasure as a sufficient criterion for robustness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.