CrossTimeBench: Benchmarking Cross-Temporal Video Reasoning in Complex Multi-Event Scenarios
Abstract
Recent advances in video reasoning have increasingly focused on improving the temporal understanding capabilities of Multimodal Large Language Models(MLLMs). However, existing benchmarks mainly focus on short clips and single-segment reasoning. While such benchmarks are useful for assessing basic temporal understanding, they are insufficient for evaluating complex multi-event scenarios where relevant evidence is distributed across different parts of a video. In this paper, we introduce the CrossTimeBench, a cross-temporal video reasoning benchmark designed to systematically evaluate model reasoning over complex multi-event videos. Each question in the benchmark requires the integration of information from more than two temporal segments, constructed through both mining and synthesis from existing datasets. The benchmark covers 8 major categories and 35 subcategories, including 2,607 videos and 8,496 QA pairs (with 6,628 unique samples due to multi-label annotations). Among them, 48.4% of open-ended QA pairs and 94.6% of closed-form QA pairs are newly constructed. In addition, we propose a fine-grained evaluation framework based on the benchmark and conduct extensive baseline experiments on a wide range of mainstream open-source and closed-source MLLMs. Our experimental results show that, although current models achieve strong performance on existing benchmarks, they still exhibit substantial limitations in cross-temporal reasoning tasks. Meanwhile, the evaluation results demonstrate noticeable differences across open-ended and closed-form QA settings, as well as among different question types. These observations provide new insights into the limitations of current video understanding models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.