acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing and Advancing Long-Video Understanding for Distributed Evidence Composition

Abstract

Multimodal large language models can process increasingly long videos, but answering questions that depend on multiple temporally distributed observations remains challenging. We investigate how evidence demand, composition structure, and visual-context allocation affect long-video understanding. We introduce DECBench, a diagnostic benchmark comprising 926 videos and 1,123 open-ended questions spanning aggregation, temporal composition, entity-state continuity, and video-grounded explanatory reasoning and summarization. To quantify per-instance evidence demand, we introduce Minimum Evidence Count (MEC), which estimates the size of a minimum sufficient evidence set and enables comparisons across composition requirements at comparable evidence-demand levels. Our diagnostic analyses reveal two complementary bottlenecks: models often fail to acquire the necessary distributed evidence, while substantial errors persist even when relevant evidence is directly provided, indicating additional difficulty in composing the available observations. We further identify a task-dependent trade-off between temporal coverage and per-observation visual fidelity. Motivated by these findings, we propose SEEK, a training-free, query-conditioned framework that selects complementary observations according to query-implied evidence requirements and adaptively allocates visual capacity under a fixed budget. Experiments across three video MLLMs and three benchmarks show consistent improvements over query-aware frame-selection baselines. Together, DECBench and SEEK provide tools for diagnosing distributed-evidence failures and improving evidence acquisition in long-video understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.