acceptodds
Under review as a conference paper at ICLR 2027

Shaped SAEs: Expanding the Toolkit for Mechanistic Interpretability

Abstract

Sparse autoencoders (SAEs) have become a widely used tool for decomposing the dense activations of large language models (LLMs) into sparse, potentially interpretable features using a learnt overcomplete dictionary. In this work, we introduce two complementary methods for shaping SAEs to enhance the semantic selectivity of the learnt dictionary atoms. The first, termed J-lens SAE, exploits the recently introduced Jacobian lens (J-lens), learning separate SAEs for the stronger and weaker eigenmodes of the averaged Jacobian between the current layer and the output layer. We expect the SAE learnt for the stronger J-lens modes to be more semantically selective, since they are more predictive of the LLM output. The second method, termed MAE-SAE, passes activations through a nonlinear masked autoencoder (MAE). The idea is that the MAE bottleneck layer will be forced to encode the relationships between the coordinates of the original token embedding, so that an SAE learnt for the MAE bottleneck layer should be more semantically selective than a standard SAE. We demonstrate that each shaping method outperforms standard SAEs on quantitative benchmarks such as SAE Bench, and discuss their qualitative attributes. Further experimentation is needed on how to combine these methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.