acceptodds
Under review as a conference paper at ICLR 2027

OBToM: Ordinal Bayesian Theory of Mind from Likelihood Judgments of Language Models

Abstract

Language models often answer theory-of-mind questions inconsistently when asked to reason directly about mental states. Methods that embed a language model within Bayesian inverse planning instead require verbalized probabilities, which are poorly calibrated, or token log-probabilities, which many deployed models do not expose. We propose Ordinal Bayesian Theory of Mind (OBToM), which separates likelihood judgment from probability estimation. A story is parsed into discrete variables on a fixed mental-state graph linking prior belief and perception to belief, and belief and goal to action. The language model is asked only ordinal questions: given two hypothetical conditions and outcomes, it judges whether one conditional likelihood is somewhat or much larger, whether the two are approximately equal, or whether the story provides no basis for judgment. An ordered-logit model converts these judgments into conditional probability tables with a shared low-rank parameterization, so a few fixed comparisons per story also constrain table rows that were never queried. Belief revision is gated by perceptual access: an agent who missed a change keeps the earlier belief. Mental states are then inferred through exact Bayesian inference over the graph, using a maximum a posteriori estimate of the tables or Hamiltonian Monte Carlo samples. We compare OBToM with direct prompting, perspective-taking, timeline, thought-tracing, and verbalized- and token-probability inverse-planning methods on the same backbone. With either estimator, OBToM is the most accurate on BigToM, FANToM, ExploreToM, and first- and second-order ToMi, and significantly so against every baseline on BigToM, ToMi, and ExploreToM. It uses at most five language-model requests per question, against 24.5 to 177.3 for thought-tracing and automated inverse planning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.