When Coordinates Change the Features: A Symmetry Test for Sparse Autoencoders
Abstract
An orthogonal change of coordinates preserves the information in language-model activations, but it can change the features learned by a sparse autoencoder (SAE). We study this dependence with a coupled symmetry test. For TopK SAEs, a rotation of the inputs has an exact counterpart in parameter space that leaves sparse codes and reconstruction loss unchanged. We initialize two SAEs at corresponding parameters and train them on paired minibatches to test whether this correspondence survives training. With Adam, the resulting dictionaries differ substantially, even after latent matching, although their reconstruction quality is similar. The discrepancy begins with Adam's coordinatewise second-moment normalization. We derive a symmetry criterion for adaptive updates and show that group-RMS updates preserve the correspondence in exact arithmetic. In our experiments, these updates reduce dictionary disagreement at comparable reconstruction quality. Thus, differences between SAE dictionaries can come from the coordinates used during training even when the language model and its activation data are fixed. Reconstruction metrics alone do not reveal this dependence. The coupled test measures it against a known correspondence and identifies optimizer choices that respect that correspondence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.