From Words to Worlds: Beyond Current Evaluation Focus, MLLMs Fail at Higher-Order Theory of Mind
Abstract
Theory of Mind (ToM), the ability to infer others' mental states from observable behavior, is fundamental to developing machines with human-level social intelligence. Despite the extensive evaluation dimensions for Multimodal Large Language Models (MLLMs), their ToM capabilities have not received sufficient attention. A comprehensive survey of ToM benchmarks reveals that while text-based ToM benchmarks for evaluating LLMs are abundant, multimodal ToM benchmarks remain scarce. Moreover, existing multimodal ToM benchmarks fail to incorporate classic ToM task paradigms from cognitive science, higher-order ToM, or visual observations better aligned with the real world. Compared to the ambiguous expressions often present in text-based ToM benchmarks, visual information is intuitive, unambiguous, and better aligned with embodied applications. In this paper, we present From Words to Worlds (FWW), a novel pipeline that enables the generation of videos from multi-scene script templates, constructing a video-based ToM benchmark named FWW-V. We further propose a rubric-based evaluation framework, FWW-EVAL, to systematically assess 14 advanced MLLMs on FWW-V. Our results reveal that current models perform poorly on ToM tasks, with a striking disparity relative to human performance. Notably, for certain question types, the majority of MLLMs fail entirely, achieving 0% accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.