acceptodds
Under review as a conference paper at ICLR 2027

Do Features in a Sparse Autoencoder Really Represent Clear Semantics?

Abstract

The sparse autoencoder (SAE) is a widely used mechanistic interpretability method that transforms dense LLM activations into sparse features. In this paper, we propose using interactions as a numerically verifiable metric to evaluate the semantic clarity of SAE features from four perspectives. (1) We observe that 19.2%–35.6% of the sampled SAE feature activations encode non-sparse interaction patterns, suggesting more diverse semantics. The average number of interactions per feature varies dramatically, and we find no clear correlation between the interaction number and the feature's activation frequency. (2) We find that SAE feature activations that encode fewer interactions tend to model simpler and more human-interpretable interactions. (3) 25.2%–41.9% of the feature activations encode interactions with severe positive-negative cancelling effects, indicating potentially less coherent semantics. (4) We find that 6.2%–19.5% of the SAE feature activations capture similar semantics across different contexts, and these activations also encode interactions with less positive-negative cancellation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.