acceptodds
Under review as a conference paper at ICLR 2027

VideoCogQA: A Controllable Benchmark for Cognitive Abilities in Video-Language Models

Abstract

Recent advances in Large Video-Language Models (LVLMs) have demonstrated promising progress in multimodal video understanding. Yet, it remains unclear whether these models possess the cognition-inspired capabilities required for higher-level reasoning, particularly when tasks involve symbolic manipulation and abstract reasoning. Existing benchmarks are predominantly built on annotated real-world videos, offering limited control over content composition and task difficulty and, consequently, limited diagnostic power for isolating specific cognitive abilities. To address this gap, we introduce VideoCogQA, a scalable and fully controllable benchmark inspired by game-based environments for systematically evaluating the cognitive capabilities of LVLMs. VideoCogQA leverages a programmatic video-generation engine that enables precise control over visual elements, temporal dynamics, and task complexity, thereby disentangling cognitive reasoning from reliance on prior semantic knowledge. The benchmark comprises tasks involving abstract concepts, symbolic representations, and multimodal information integration, with difficulty systematically varied through Python-based game scenarios. Extensive experiments reveal substantial limitations in current LVLMs: even state-of-the-art models such as Qwen2.5-VL-72B achieve only 48.8% average accuracy on tasks involving abstract concepts, while performance further drops by 15% as task complexity increases. These results suggest that, despite strong performance on existing video understanding benchmarks, current LVLMs still struggle with robust and systematic cognitive reasoning under controlled increases in task complexity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.