Belief Without Behavior: Evaluating Social Reasoning-to-Action in Multimodal LLM Agents
Abstract
Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under varied ToM constraints. Evaluating 16 models, including 14 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level diagnostics identify two recurring bottlenecks: weak directional signal production, and limited correspondence between available signals and receiver choices. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, is the only evaluated system that clears both bottlenecks under the canonical experimental geometry, but its deceptive strategy collapses under non-canonical spatial layouts, indicating that explicit belief-action coupling is a promising, though not on its own sufficient or robust, architectural ingredient for this class of tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.