acceptodds
Under review as a conference paper at ICLR 2027

HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

Abstract

Recent research often evaluates models’ video understanding abilities through tasks such as summarization, temporal reasoning, and cross-modal reasoning. Existing benchmarks typically focus on a single task and a single level of granularity for each video, limiting their ability to provide a comprehensive evaluation of video understanding. We introduce HAVEN, a benchmark that links visual and textual annotations at the frame, shot, and video levels. Built from six existing video-summarization datasets, HAVEN contains 11,055 videos and uses a 365-video subset with human-audited textual annotations for standard evaluation. Its task suite covers summarization, temporal ordering, cross-modal grounding, and saliency ranking. Empirical experiment results reveal a comprehensive and fine-grained evaluation of the video understanding capabilities of widely used vision-language models. These observations show that textual-summary scores alone do not capture the variation measured by visual selection and alignment tasks. HAVEN provides a shared annotation structure for examining these complementary aspects of video understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.