HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
Abstract
Recent research often evaluates models’ video understanding abilities through tasks such as summarization, temporal reasoning, and cross-modal reasoning. Existing benchmarks typically focus on a single task and a single level of granularity for each video, limiting their ability to provide a comprehensive evaluation of video understanding. We introduce HAVEN, a benchmark that links visual and textual annotations at the frame, shot, and video levels. Built from six existing video-summarization datasets, HAVEN contains 11,055 videos and uses a 365-video subset with human-audited textual annotations for standard evaluation. Its task suite covers summarization, temporal ordering, cross-modal grounding, and saliency ranking. Empirical experiment results reveal a comprehensive and fine-grained evaluation of the video understanding capabilities of widely used vision-language models. These observations show that textual-summary scores alone do not capture the variation measured by visual selection and alignment tasks. HAVEN provides a shared annotation structure for examining these complementary aspects of video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.