acceptodds
Under review as a conference paper at ICLR 2027

Receptors of Language Models: From Singular Directions to Token Predictions

Abstract

While mechanistic interpretability has largely investigated the information/context flowing through the transformer architecture, the structural blueprint statically/inherently encoded in its parameters remains comparatively underexplored, leaving open a more fundamental question, what predictive structure is already encoded in the weights (inherent to model parameters)? We show that the static weight matrices of transformer components contain sparse, token-aligned directions, which we term , that show a strong causal link with model predictions. This structure is consistent across a wide family of architectures. Across 21 open-weight models spanning 5 architecture families (, , , , and ), we demonstrate the presence of receptors providing an interpretability lens that is applicable to wider range of contexts. These receptors emerge and stabilize over training. Additionally, we find that a kurtosis-based analysis of individual receptors reveals additional structure in their weight-space geometry. Our work provides a detailed examination of the model's inherent structural properties, which could serve as a complementary foundation for studying, measuring, and intervening in learned behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.