Keep in mind: Can LLMs imagine object relations without speaking?
Abstract
Humans can imagine relations among objects without verbalizing them, yet visual working memory holds only a handful of objects—about three to four—without verbal support. It remains poorly understood whether large language models (LLMs) possess this ability and what its limits are. The question matters for downstream applications: reliable spatial reasoning about object layouts is essential for embodied AI, and whether a model can reason in its hidden space affects the effectiveness of chain-of-thought (CoT) monitoring in AI safety. To characterize this capability, we introduce a simple evaluation setting that separates three requirements: thought-privacy, self-consistency, and cross-order agreement. We evaluate 13 models spanning a wide size range (1B to 744B parameters), from standard transformers to looped transformers, with and without CoT. Our experiments show that, for current open-source reasoning models, thought-privacy is hard to achieve regardless of model size or architecture. When thought-privacy is achieved, self-consistency and cross-order agreement are hard to attain. These results indicate that explicit text continues to play an important role in sustaining relational consistency. To examine the internal activity accompanying these answers, we extend the Jacobian lens (J-lens) to measure how hidden states align with groups of object names and direction words in Qwen3.8-27B. Together with the relation-readout and reasoning-text replay results, these findings suggest that relational consistency across questions relies on explicit records of prior reasoning in the context rather than on a stable, unspoken private state. Overall, models struggle to maintain a fixed arrangement of object relations without verbalizing it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.