acceptodds
Under review as a conference paper at ICLR 2027

Iconic Format Costs Language Models Accuracy: A Preregistered Three-Format Study

Abstract

It is a common claim in the cognitive sciences that mental representations come in different formats. A representational format is a way of information organization that dictates a representation's specific computational profile. For example, discursive representations—notably, natural language representations—are decomposable into constituent representations that combine in certain canonical ways. Iconic representations, meanwhile, are representations that represent an object by way of mirroring its arrangement. A typical example of an iconic representation is that of a typical map, which mirrors, in miniature, the actual distance relations between the places it depicts. Whether language models build internal models of spatial structure is disputed and whether any such models are iconically formatted has rarely been asked. This set of pre-registered studies aims to make progress on this question via a multi-pronged strategy of evaluating behavior and model mechanisms. Study 1 evaluates three open-weight and two frontier models on informationally equivalent but differently formatted reasoning tasks. Study 3, meanwhile, subjects these models to the same tasks but removes their written-out reasoning, asking for the answer alone. More specifically, these studies ask the language model to supply correct answers to natural language sentences, diagrams, and "tabular" formats. These studies find that the open-weight models perform worse on iconic tasks compared to discursive or tabular tasks, and by a wide margin (−0.92 log-odds against tables in Study 1; −25.1 percentage points in the graph family, and −6.0 to −13.0 percentage points in the other three, depending on the scorer), suggesting that contemporary language models struggle to reason over icons. The frontier models are at or near ceiling in all three formats outside the graph family, so outside it these tasks are uninformative about them; within it, they performed worse on the iconically formatted questions (5.2 and 6.0 percentage points). Removing written-out reasoning narrows the gap (−0.40 log-odds). Because studies 1 and 3 are behavioral, they cannot tell us about whether language models use, or even have, iconic representations. Study 2 endeavored to test this by probing the residual streams of the model while performing these tasks. The outcome of the study was negative, insofar as the results were entirely determined by choice of readout, but supplied an important methodological result. That result is that linear decodability from the residual stream cannot discriminate between a language model's having iconic representations versus a discursive representations of spatial facts. Three conclusions follow from this work. First, there is some evidence that current language models struggle to reason over icons. Second, a common method for studying the internal representations of language models, linear probing, cannot settle questions about format in particular. Third, these results point to better designs for such studies: controls that detect when a readout is blind to format, and causal tests of whether a model uses the information it encodes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.