A Testable Theory of Atomic Features
Abstract
We develop and test a theory of language model representations in which there exist "atomic" features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of atomic features most prevalent in the training distribution. This "recovery principle" yields three testable predictions: many features in small SAEs are shared by all larger SAEs, SAEs trained on different distributions share features prevalent in both, and sufficiently large SAEs recover both parent and child features. In contrast to conventional wisdom that SAEs features are unstable and "split" as size increases, we find that these predictions hold on SAEs of sizes ranging from 512 to 131,072 trained on two large embedding models. From a theoretical perspective, our results suggest the promise of a scientific theory of representations based on atomic features. Practically, our results suggest the promise of scaling SAEs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.