acceptodds
Under review as a conference paper at ICLR 2027

Can You Hear My Mind? Catching Theory of Mind in Voice and Pragmatics

Abstract

The same words can convey different intentions depending on how they are spoken. Interpreting these cues alongside verbal content draws on Theory of Mind (ToM), the capacity to perceive and infer others’ mental states. Yet existing ToM benchmarks primarily assess mental-state reasoning from textual or visual events, while audio reserach largely focuses on state recognition rather than cognition. This leaves unclear how vocal delivery and pragmatics shape mind. We introduce Overtone, the first multimodal ToM benchmark to systematically assess how models use vocal and pragmatic cues to infer mental states. Overtone comprises 2,335 social-interaction scenarios with human-recorded speech, enables comprehensive, fine-grained evaluation across diverse classical ToM tasks and a broad spectrum of ToM abilities. A human study and evaluations of 16 speech and omni-modal models reveal that models struggle to maintain coherent mental-state representations and integrate vocal cues into ToM reasoning. Hence, we introduce PEAR (Prediction-Error-Adaptive Reasoning), a training-free ToM reasoning framework inspired by predictive coding in the human brain. Mental-state hypotheses guide the interpretation of multimodal cues, while metacognitive regulation the reasoning process. Extensive experiments show that PEAR improves performance across backbones, outperforms several reasoning baselines, and generalizes across diverse ToM datasets. Our findings highlight current multimodal models' limitations in inferring mental states from vocal and pragmatic cues. Together, our work provides a benchmark and reasoning framework for assessing and improving this capability, advancing multimodal social reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.