Shared Conditional Transformations of Arithmetic Directions in Language Models
Abstract
Mechanistic interpretability seeks to explain model behavior, in part by identifying feature directions in internal representations and testing their effects through interventions. Yet a central question is whether shared structural principles can explain and predict how representations change across contexts. In controlled addition, we find that arithmetic directions exhibit condition-dependent variation described by low-dimensional orthogonal transformations shared across values at certain layers. These transformations predict target-condition directions for values excluded from fitting. At the reporting layer of each of four language models, hidden-state interventions using directions predicted by these shared transformations achieve exact complete target-answer match rates of 71.13%–99.43%, compared with 29.66%–41.68% using untransformed source directions. In three models, these effects approach those of directions estimated directly under the target condition. The shared-relation pattern also recurs after independent refitting in a tens-column setting. We further characterize a more constrained structure in terms of direction-level approximate conditional equivariance: powers of a shared orthogonal operator approximate the transformations associated with cyclic condition shifts. Predictions under this description induce expected output changes at multiple tested layers. Together, these findings suggest that interpretable structure can lie not only in feature directions themselves, but also in the shared, predictable relations organizing their variation across contexts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.