acceptodds
Under review as a conference paper at ICLR 2027

When Do Sparse Autoencoders Recover Features, and What Do They Learn Instead?

Abstract

Sparse autoencoders (SAEs) reconstruct activations as nonnegative sums of learned directions, and each latent's activation is read as the strength of one feature of the data. Training sees only reconstruction and sparsity, never the features, and learned latents are known to split, absorb and merge features and to change between runs. One explanation is limited inference: a one-layer encoder can provably fall short of the best sparse code chosen separately for each input, and more expressive encoders have been reported to recover such codes better and to yield more interpretable latents. We remove that limitation, building synthetic data whose features are identifiable and whose strengths a one-layer ReLU encoder computes exactly, and vary only how the strengths of co-occurring features depend on each other while directions, frequencies, marginal strength laws and covariance stay fixed. For independent exponential strengths, every sufficiently near-optimal affine-ReLU SAE of arbitrary finite width recovers each feature's additive contribution through fixed groups of latents as the penalty vanishes, whereas every such near-optimum retains nonvanishing contribution error when co-occurring strengths are often nearly equal. Adding a second encoder layer makes a cheaper non-recovering code available even under independence, while penalizing the number of active latents with one-layer responses restores individual recovery as its price vanishes. On a 16-dimensional synthetic source, SAEs trained on the coupled data reach a lower loss than any code on the generating directions. Recovery failures therefore need not come from limited inference, and when inference is already exact, a more expressive encoder can cause them.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.