acceptodds
Under review as a conference paper at ICLR 2027

Dissecting Theory of Mind in Large Language Models via Bayesian Modeling

Abstract

Large Language Models (LLMs) perform well on narrative Theory of Mind (ToM) benchmarks, but whether they can jointly infer an agent's desires and beliefs from a sequence of spatial actions remains largely untested. We evaluate ten LLMs on a dynamic spatial ToM task in a tripartite comparison against archived human ratings and a normative Bayesian Theory of Mind (BToM) model, and use a family of alternative observer models—each removing or distorting one component of BToM—as a diagnostic toolkit for LLM inference. LLMs varied substantially. The best models correlated with human judgments at levels approaching, but not exceeding in both desire and belief, the human–BToM alignment. The desire judgments of lower-performing models were best captured by a variant that ignores action costs, and their belief judgments by heterogeneous suboptimal variants. We then asked models to reconstruct, at every step of the trajectory, the agent's initial belief using only the evidence up to that step while the full trajectory remained in context (Every-step condition). In this format, the belief judgments of even the best models no longer correlated with human judgments and were best accounted for by a new Hindsight model that projects the later-observed world state onto the agent's earlier belief, a pattern reminiscent of the human hindsight bias and curse of knowledge. These results suggest that successful final-state social reasoning in LLMs does not imply consistent mental-state inference over time, and that a Bayesian inverse-planning framework allows such failures to be decomposed into specific components and tested against alternative models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.