How to track your tangent space: the geometry of the linear representation hypothesis
Abstract
We examine how local geometry shapes interpretability by studying the tangent structure of the activation manifold of language models: we show that hidden states can be decomposed into linear and multidimensional components by projection onto the tangent space. Using this decomposition, we then show that linear components dominate the interpretable signature of the network. We then perform a similar decomposition of the recently proposed -lens, to concentrate its on-manifold versus off-manifold data. We give a full analytic description of a tangent-native -lens, and show that the normal contribution to the -lens depends on the off-manifold extension of the network: in particular, the workspace readout can be changed arbitrarily subject to mild technical constraints, which could prevent a misaligned network from being detected. Empirically, we show that there is not a simple fix for this: the on-manifold Jacobian shows substantial cancellation in mid-to-late layers, and the -lens outputs materially rely on their off-manifold projection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.