Principled Nonlinear Concept Erasure and Steering
Abstract
Learned representations often encode sensitive attributes, such as gender or race, that may be undesirable or unsafe to retain. Concept erasure aims to remove this information while preserving task-relevant content. Existing methods remain incomplete: linear methods provide guarantees only against linear probes, while nonlinear methods achieve stronger empirical erasure but lack formal guarantees and often sacrifice utility. In this work, we introduce a principled conditional autoencoder framework for nonlinear concept erasure. We follow the principle that erasure should remove only the sensitive attribute and nothing else, and we show that, under the stated assumptions, perfect reconstruction together with population-level independence is *sufficient* for perfect nonlinear guardedness, with preservation guarantees in the imperfect-reconstruction regime. Empirically, we attain a superior utility-erasure trade-off across multiple datasets, sensitive attributes, and language models, surpassing even the supervised variants of competing methods despite using no target-task labels, while being simpler to optimize and converging substantially faster. Beyond erasure, our framework yields *concept conversion (representation steering)* as a by-product: by changing only the target concept at inference time, it produces representations corresponding to different concept values—at no additional training cost—a capability no prior erasure method provides, and on which we likewise demonstrate strong empirical performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.