Symmetry, Not Dimension: When Sparse-Coding Atoms in a Feature Plane Are Identifiable
Abstract
Sparse autoencoders trained from different seeds disagree, which concurrent work attributes to feature manifolds of dimension ; in the sharpest analysis, from an isotropy assumption. For a feature plane, the sparse-coding objective under a rigid rotation of its atoms is a circular cross-correlation between the angular density and a configuration-dependent cost kernel, flat along the rotation iff the density's relevant Fourier coefficients vanish, else pinned with computable curvature. Under , dimension is necessary for that flatness, not sufficient. On densities from models' own rings, twelve-seed dictionaries meet predictions the density condition makes and symmetry cannot: agreement at an installed harmonic (strongest in of Gemma-2-2B tilts at no symmetry explains; chance ), and its loss when one Fourier coefficient is removed from a spread density, including the number circles of Llama-3.1-8B and Pythia-6.9B, but not from clustered rings. Only ReLU+ seeds reliably sit near their landscape's minimum; TopK's in-plane optimum at is flat on full-support densities. On released dictionaries and our own, ablating a dictionary's plane for a concept transfers across dictionaries in all cells, its most selective latent in , failing where dictionaries disagree on its item. Making the data exactly symmetric in one plane removes the seeds' agreement on its atoms (matched atoms apart) and reinstalling one harmonic restores it; a plane found in natural text without probes transfers too. Claims should attach to the plane, not the atom. Code, records and some trained dictionaries are in the supplement; all will be public at https://github.com/xxx/xxx upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.