acceptodds
Under review as a conference paper at ICLR 2027

Do LLMs Learn Universal Symbolic Representations? An Early Exploration

Abstract

Exploring internal cognitive patterns of large language models (LLMs) has attracted growing research interest. In this paper, we use interaction patterns as an interpretable and verifiable metric to rigorously decompose the inference logic of an LLM into symbolic AND-OR interactions, with theoretical guarantees of explanatory completeness. We find that mainstream open-weight LLMs exhibit highly consistent symbolic interaction patterns despite substantial differences in parameter scale and training data. Compared with model-specific, non-shared interactions, these cross-model shared interaction patterns exhibit lower structural complexity and weaker positive-negative cancellation, making them more compelling as inference patterns. We further show that a large fraction of LLM prediction scores is attributable to these shared interaction patterns. Overall, our work validates the hypothesis that LLMs encode inherent symbolic patterns and provides a precise and verifiable decomposition of cognitive patterns that are universally shared across LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.