The Shape of Depth: Consistent Decoder Geometry in Sparse Autoencoders Across LLM Layers
Abstract
To disentangle the superposed semantic units inside large language models (LLMs), sparse autoencoders (SAEs) have become a go-to choice. Yet, they remain layer-wise snapshots that are blind to how representations evolve throughout the full forward pass to the next token. In this work, we interrogate the geometrical structure of SAE decoders across layers of GPT-2 and find a kernel of information that undergoes a consistent rotation from layer to layer – perhaps a form of LLM working memory. We thus show that the top activation directions of different layers exhibit a clear rotational trend in cosine similarity, and canonical correlation analysis of decoder structure further reveals that residual activation dimensions map into a persistent structure. To show practical benefit, we construct a plane spanned by the leading activation directions of two layers via Gram–Schmidt orthogonalization and derive a rotation operator on this plane. We control this rotation to applied activation-concept spaces during cross-layer transfer learning, which improves SAE training, and yields fewer dead SAE features by treating rotation angle as a tunable hyperparameter. These insights suggest that transformations of a geometric object persistent across layers can be monitored and leveraged, such as to enhance SAE training and performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.