IVU-Bench: Benchmarking Multimodal Large Language Models for Video Understanding under Incomplete Observations
Abstract
Multimodal large language models (MLLMs) have demonstrated strong video understanding, yet existing evaluations largely focus on recognizing explicitly observable content. Real-world video understanding often requires inference beyond observation, including predicting missing events, recovering spatial relations from partial evidence, and reasoning under modified or hypothetical conditions. We introduce IVU-Bench, a benchmark for evaluating MLLMs under incomplete observations, comprising 2,032 question-answer pairs, and 13 tasks across temporal, spatial, and factual dimensions. Temporal incompleteness involves inferring missing events or temporal structure; spatial incompleteness involves recovering object locations and relations from partial observations across time and viewpoints; and factual incompleteness involves reasoning about event dynamics under modified or hypothetical conditions absent from the original video. Questions from multiple task types are grounded in shared videos, enabling controlled cross-task comparisons while reducing content-related confounds. To assess whether correct answers are supported by appropriate visual evidence, IVU-Bench provides object-level spatiotemporal annotations for target evidence and semantically related or irrelevant distractors. We further evaluate models with three complementary metrics: Answer Accuracy, Evidence Discrimination, and Joint Answer–Evidence Score, measuring answer correctness, target-evidence preference, and answer–evidence consistency, respectively. Experiments across representative MLLMs reveal substantial weaknesses across all three dimensions, especially in spatial reasoning from partial visual evidence, highlighting the gap between recognizing visible content and inferring unobserved information.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.