acceptodds
Under review as a conference paper at ICLR 2027

The Shape of Depth: Consistent Decoder Geometry in Sparse Autoencoders Across LLM Layers

Abstract

To disentangle the superposed semantic units inside large language models (LLMs), sparse autoencoders (SAEs) have become a go-to choice. Yet, they remain layer-wise snapshots that are blind to how representations evolve throughout the full forward pass to the next token. In this work, we interrogate the geometrical structure of SAE decoders across layers of GPT-2 and find a kernel of information that undergoes a consistent rotation from layer to layer – perhaps a form of LLM working memory. We thus show that the top activation directions of different layers exhibit a clear rotational trend in cosine similarity, and canonical correlation analysis of decoder structure further reveals that residual activation dimensions map into a persistent structure. To show practical benefit, we construct a plane spanned by the leading activation directions of two layers via Gram–Schmidt orthogonalization and derive a rotation operator on this plane. We control this rotation to applied activation-concept spaces during cross-layer transfer learning, which improves SAE training, and yields fewer dead SAE features by treating rotation angle as a tunable hyperparameter. These insights suggest that transformations of a geometric object persistent across layers can be monitored and leveraged, such as to enhance SAE training and performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.