Receptors of Language Models: From Singular Directions to Token Predictions
Abstract
While mechanistic interpretability has largely investigated the information/context flowing through the transformer architecture, the structural blueprint statically/inherently encoded in its parameters remains comparatively underexplored, leaving open a more fundamental question, what predictive structure is already encoded in the weights (inherent to model parameters)? We show that the static weight matrices of transformer components contain sparse, token-aligned directions, which we term , that show a strong causal link with model predictions. This structure is consistent across a wide family of architectures. Across 21 open-weight models spanning 5 architecture families (, , , , and ), we demonstrate the presence of receptors providing an interpretability lens that is applicable to wider range of contexts. These receptors emerge and stabilize over training. Additionally, we find that a kurtosis-based analysis of individual receptors reveals additional structure in their weight-space geometry. Our work provides a detailed examination of the model's inherent structural properties, which could serve as a complementary foundation for studying, measuring, and intervening in learned behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.