Mind-VBench: A Diagnostic Benchmark for BDI-Consistent Behavior in Video Generation
Abstract
Recent video generators produce increasingly realistic and coherent videos, yet visually plausible videos may still depict characters acting inconsistently with the information available to them or with the goals they pursue. We study this discrepancy by evaluating whether generated character behavior remains consistent with the mental states established by a scenario. To make this question systematically evaluable in video generation, we introduce Mind-VBench, a diagnostic benchmark that operationalizes the Theory of Mind (ToM) question as observable behavioral constraints, using the Belief–Desire–Intention (BDI) framework from cognitive science and agent modeling. It contains 360 instances organized into three BDI components, six diagnostic dimensions, and twelve tasks. We evaluate representative video generation models under text-to-video and image-to-video settings. Current models frequently fail both to establish the situation required for behavioral judgment and to generate character behavior consistent with the specified BDI constraints. These findings show that current video generators do not reliably maintain mental-state consistency in character behavior, highlighting a limitation for human-centered world simulation. Our demo is available in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.