acceptodds
Under review as a conference paper at ICLR 2027

De-Absorption Reparameterization: Post-hoc Repair of Feature Absorption in Sparse Autoencoders.

Abstract

Sparse autoencoders (SAEs) are widely used to decompose neural-network activations into sparse, human-interpretable features. However, when the underlying features are hierarchically related, the same sparsity that promotes interpretability can distort their representation. In particular, a broad parent feature may stop activating when one of its more specific child features fires. This phenomenon, known as feature absorption, makes individual SAE features less reliable for interpretation and intervention. Existing approaches mainly aim to reduce absorption by modifying SAE training. In contrast, we ask whether feature absorption can be corrected after the SAE has already been trained. To address this question, we introduce De-Absorption Reparameterization (DAR), a post-hoc method that corrects feature absorption for a supplied relation between a parent feature and one or more child features in a trained SAE. DAR redistributes absorbed parent-related information from the child decoder back to the parent feature, while applying a compensating decoder update that preserves the SAE reconstruction exactly. Since reconstruction is preserved for any transfer coefficient, reconstruction alone cannot identify the correct repair. DAR therefore estimates the coefficients from parent–child activation patterns using sign-constrained least squares, ensuring that the parent activation only increases while all latent activations remain nonnegative. We evaluate DAR on JumpReLU and BatchTopK SAEs for Gemma-2-2B and CLIP ViT-L. Given the supplied parent–child relations, DAR restores parent activations on absorbed inputs across both language and vision. On the main language benchmark, DAR restores parent activation on up to 44.6% of previously absorbed examples. The repaired activations are more usable for their intended interpretation: intended-latent AUROC on absorbed inputs increases from 0.535 to 0.941. Casually, ablating the repaired parent produces a 15× larger effect on the relevant model logit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.