De-Absorption Reparameterization: Post-hoc Repair of Feature Absorption in Sparse Autoencoders.
Abstract
Sparse autoencoders (SAEs) are widely used to decompose neural-network activations into sparse, human-interpretable features. However, when the underlying features are hierarchically related, the same sparsity that promotes interpretability can distort their representation. In particular, a broad parent feature may stop activating when one of its more specific child features fires. This phenomenon, known as feature absorption, makes individual SAE features less reliable for interpretation and intervention. Existing approaches mainly aim to reduce absorption by modifying SAE training. In contrast, we ask whether feature absorption can be corrected after the SAE has already been trained. To address this question, we introduce De-Absorption Reparameterization (DAR), a post-hoc method that corrects feature absorption for a supplied relation between a parent feature and one or more child features in a trained SAE. DAR redistributes absorbed parent-related information from the child decoder back to the parent feature, while applying a compensating decoder update that preserves the SAE reconstruction exactly. Since reconstruction is preserved for any transfer coefficient, reconstruction alone cannot identify the correct repair. DAR therefore estimates the coefficients from parent–child activation patterns using sign-constrained least squares, ensuring that the parent activation only increases while all latent activations remain nonnegative. We evaluate DAR on JumpReLU and BatchTopK SAEs for Gemma-2-2B and CLIP ViT-L. Given the supplied parent–child relations, DAR restores parent activations on absorbed inputs across both language and vision. On the main language benchmark, DAR restores parent activation on up to 44.6% of previously absorbed examples. The repaired activations are more usable for their intended interpretation: intended-latent AUROC on absorbed inputs increases from 0.535 to 0.941. Casually, ablating the repaired parent produces a 15× larger effect on the relevant model logit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.