SumReason: A Benchmark for Reasoning-Driven Query-Conditioned and Standard Video Summarization
Abstract
Video summarization benchmarks have largely evolved along two disconnected tracks: standard summarization datasets that assume a single notion of importance, and query/text-conditioned datasets whose queries often reduce to simple object or concept retrieval/matching. As a result, current benchmarks do not adequately test whether models can adapt summaries to user intent or perform the temporal, semantic, and knowledge-based reasoning required in realistic settings. We introduce SumReason, a unified benchmark for both standard and query-conditioned video summarization that explicitly targets this gap of reasoning-intensive video understanding. The dataset contains 37 videos, 60 video-query pairs, and 25 query-independent annotation sets, with 850 total annotations from 10 annotators. Its query hierarchy spans Direct Visual, Temporal Reasoning, High-Level Semantic Understanding, and External Knowledge Reasoning, and is designed to shift focus toward both salient and secondary events. SumReason combines strong annotation reliability with increased task difficulty: it achieves higher inter-annotator agreement than legacy benchmarks in the standard setting, while query-conditioned summaries overlap with standard summaries by only 30–40% on average, confirming substantial query-induced semantic shifts. Across extensive evaluations of >15 video summarization methods and frontier open- and closed-source MLLMs, current systems exhibit substantial limitations on SumReason: standard summarization methods show markedly smaller gains than on established datasets, while MLLMs reveal a persistent reasoning gap in the query-conditioned setting. We position SumReason as a benchmark for studying intent-aware summarization moving beyond surface-level visual matching.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.