acceptodds
Under review as a conference paper at ICLR 2027

Analyzing Spatial Cognition via Behavioral Interpretability

Abstract

What are LLMs learning to solve spatial cognition tasks? Interpretability techniques usually inspect a model's verbalized reasoning or internal activations, but chain-of-thought traces can be unfaithful and activations are inaccessible for closed models. Behavioral psychology faces the same two obstacles, since self-reports are unreliable and the brain is hard to observe, and instead studies the mind by modeling people's choices and mistakes. We propose to study LLMs the same way by synthesizing symbolic programs that predict their choices and mistakes. On four interactive spatial cognition tasks, our evolutionary search finds programs that match each model's action distribution to within 0.07 nats while being two orders of magnitude simpler than a memorization baseline. The programs reveal that the strongest models perform shortest-path planning in maze navigation, raster scanning and spatial memory-based pruning to find treasures in the Cambridge spatial working memory task, and lookahead planning over visual waypoints for egocentric navigation. On the other hand, weaker models perform greedy goal search in mazes and get stuck in dead ends, exhibit poor retention in spatial working memory and reopen boxes known to be empty, and oscillate in place for egocentric navigation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.