SMORE: Minimal-Pair Videos for Social, Moral and Rational Reasoning in MLLMs
Abstract
Understanding human behavior involves interpreting actions in relation to agents’ beliefs, goals and social norms. Small differences in behavior can change these interpretations even when the surrounding context remains similar. To test whether multimodal large language models (MLLMs) reliably distinguish such cases, we introduce SMORE, a minimal-pair benchmark for evaluating Social, MOral, and Rational reasoning. It comprises 427 video pairs (854 videos) and 1,281 evaluation questions spanning belief-conditioned rationality, helping and hindering and judgements of intentionality, responsibility and norm violations. Videos within each pair share the same actors, environment, and overall interaction while differing in a single socially-relevant factor that changes the target judgement. This design probes whether model judgements track relevant behavioral differences under closely matched conditions. We evaluate models under two settings. First, we have paired-video evaluation, where both videos from a minimal pair are presented jointly to measure relative discrimination performance. Second, we have single-video evaluation, where each video is evaluated independently to measure absolute performance. From the second setting, we additionally compute Joint-Pair Accuracy, which gives credit only when both videos in a pair are individually answered correctly. On SMORE, humans reach 96% Joint-Pair Accuracy while the best model, Gemini-3.5-Flash, reaches only 63.2%. These results suggest that, despite strong performance on video action understanding, current MLLMs still fall substantially short of humans in reasoning about the subtle differences in human behavior that determine social, moral, and rational judgments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.