V-Horizon: A Diagnostic Benchmark for Ultra-Long Video Understanding
Abstract
Understanding videos that span tens of hours requires not only identifying sparse evidence but also interpreting it in context. We introduce V-Horizon, a diagnostic benchmark comprising 660 six-choice questions from 34 complete TV series, with durations ranging from 4 to over 50 hours. The benchmark evaluates four complementary dimensions: Temporal Tracking (T), Narrative Understanding (N), Visual Perception (V), and Visual Localization (W), separating questions that are largely answerable from textual information from those that require visual evidence. We evaluate 15 models and systems, including variants based on frame selection and visual-token compression. Controlled input comparisons show that subtitle-only settings yield substantially larger gains on T and N than on V and W, confirming the distinct dependence of the latter on visual information. Providing oracle evidence improves W accuracy by up to 38.6 percentage points, revealing a strong sensitivity to evidence accessibility. We further introduce Agentic Skills, a retrieval-based framework that searches subtitles and dense visual captions and can optionally invoke a local VLM for targeted inspection. Agentic Skills achieves 70.3% overall accuracy, compared with 55.3% for the strongest holistic configuration in our main comparison. Even without runtime visual tools, it retains 69.6% accuracy, indicating that indexed visual descriptions and targeted retrieval account for much of the performance gain. Overall, V-Horizon provides a controlled testbed for studying evidence access in long-form video understanding while helping distinguish capability limitations from failures to retrieve the information needed to answer a question.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.