acceptodds
Under review as a conference paper at ICLR 2027

The Saussurean Representation Hypothesis: A Causal Account of Representational Structures in LLMs

Abstract

Mechanistic Interpretability represents a dominant set of approaches and often seeks to localize model responses to particular input features in specific attention heads, neurons, or circuits, typically validating such localizations by ablating the identified model components and measuring degradation of the targeted capability. This target-centric approach frequently fails to engage with a core tenet of semiotics: that meaning arises from the relations among concepts rather than the concepts themselves. Here we examine collateral damage, investigating how ablating one concept within a model affects performance on other similar concepts. We introduce a non-parametric, information-theoretic method for performing ablations. By estimating attention-head-level mutual information, we reveal which components of a model encode information related to specific features in a data distribution. Examining languages and WordNet-derived knowledge domains, we show that model representations recover independently defined similarity structure in both, placing languages of the same subfamily, and topically related domains, closer together in their representation space. Ablating heads with the highest pointwise mutual information for a given language causes significant damage to the model's performance on that language, and also affects members of the same language subfamily. When we apply this non-parametric method to 51 WordNet domains derived from Wikipedia across Qwen, Llama, OLMo, and Gemma models, we find that collateral damage across concepts consistently correlates with the semantic distance between them. We term this the Saussurean Representation Hypothesis: in LLMs, the functional representation of a language or conceptual domain is constituted in substantial part by its relations to neighboring representations. Consequently, causal localization cannot be evaluated only by what an intervention removes; what it removes in parallel with the target reveals the structure of the representation itself.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.