acceptodds
Under review as a conference paper at ICLR 2027

Absorption Aware Editing: Predicting and Closing the Leakage Channels of SAE Based Concept Erasure

Abstract

Sparse autoencoders (SAEs) provide interpretable latent features that can be used to edit concepts in language models without retraining. The usefulness of such edits depends on concept-associated latents providing reliable control over the corresponding information in the model representation. However, ablating these latents often leaves substantial concept information behind, and existing methods provide little guidance about when these edits will fail or how to repair them. We show that the extent of feature absorption, a known SAE pathology, quantitatively predicts the resulting concept leakage, and surviving information is distributed across absorbing SAE latents and the reconstruction residual. Based on this decomposition, we propose a training-free, absorption-aware editing method that removes the target concept direction from both. Across multiple SAE training methods and model families, the proposed method substantially reduces concept leakage with minimal effects on unrelated concepts and model utility. Together, these results turn a systematic source of SAE editing failure into one that can be diagnosed, predicted, and corrected.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.