Do LLMs Learn Universal Symbolic Representations? An Early Exploration
Abstract
Exploring internal cognitive patterns of large language models (LLMs) has attracted growing research interest. In this paper, we use interaction patterns as an interpretable and verifiable metric to rigorously decompose the inference logic of an LLM into symbolic AND-OR interactions, with theoretical guarantees of explanatory completeness. We find that mainstream open-weight LLMs exhibit highly consistent symbolic interaction patterns despite substantial differences in parameter scale and training data. Compared with model-specific, non-shared interactions, these cross-model shared interaction patterns exhibit lower structural complexity and weaker positive-negative cancellation, making them more compelling as inference patterns. We further show that a large fraction of LLM prediction scores is attributable to these shared interaction patterns. Overall, our work validates the hypothesis that LLMs encode inherent symbolic patterns and provides a precise and verifiable decomposition of cognitive patterns that are universally shared across LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.