acceptodds
Under review as a conference paper at ICLR 2027

Beyond Coordinatewise Inference in Sparse Autoencoders

Abstract

Sparse autoencoders (SAEs) are often used as if each learned coefficient represented an independent, editable concept. We show this assumption is structurally false. Standard SAE encoders, including affine, ReLU, TopK, and Gated SAEs, compute each coefficient from a single affine score of the activation. The selection step in TopK compares coordinates, but each score remains a projection of the input alone. We call these coordinatewise-score encoders. Because SAE dictionaries are overcomplete, such scores cannot each respond to their own feature's decoder direction while ignoring all others, so coefficient errors persist even when the encoder is given the dictionary that generated the data. The errors have two consequences: (1) after a single-coordinate edit is decoded into activation space and the resulting activation is re-encoded, the inferred code change can involve unedited coordinates or even move the target coordinate in the wrong direction, creating attribution ambiguity. (2) During training, the same limitation lets the decoder adapt to compensate for the coefficient error instead of the encoder correcting it, which we call decoder absorption, a mechanism of which feature absorption—a well-known SAE phenomenon—is a special case. We show that a cross-coordinate encoder can overcome this failure, and introduce a local consistency loss that further improves edit consistency and reconstruction quality over standard affine and scalar-nonlinear baselines in the large majority of configurations tested.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.