CogViD: A Comprehensive Cognitive Benchmark for Evaluating Video MLLMs
Abstract
Video understanding benchmarks have evolved from short clips to hour-long videos and from single to dozens of tasks, pushing multimodal large language models (MLLMs) toward long-form, multi-task comprehension. While existing benchmarks provide reasonable coverage of static perception, object permanence and episodic-like memory receive only partial coverage and often lack high-level planning via dedicated sub-tasks. In this paper, we propose CogViD, a cognition-inspired dataset that introduces new tasks and provides a comprehensive benchmark organized into hierarchical levels. CogViD contains QA pairs over videos and hours of footage, spanning sub-tasks and questions in both multiple-choice and open-ended form. Furthermore, we propose a template-driven multi-expert annotation pipeline in which level-specific templates drive QA generation, and verification is assigned to separate models to mitigate the bias of a model validating its own outputs, and blind-answerable and near-duplicate answers are filtered before human review. Our comprehensive evaluation across state-of-the-art MLLMs reveals a large and consistent gap in maintaining quantities and states over time (including counting over different frames), procedural memory to identify missing steps, and open-ended goal and error reasoning. CogViD not only offers a robust benchmark but also paves the way for future research toward human-like intelligence in video understanding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.