Beyond Frame Selection: Decoupled Adaptive Evidence Selection for Video Understanding with MLLMs
Abstract
In video understanding, MLLMs often need to reason over visual content whose relevance to a question varies over time. Existing frame selection and visual-token pruning methods typically rely on a unified relevance ranking or fixed selection granularity, overlooking that visual-text and ordinary visual content can exhibit different relevance distributions and contribute differently to different questions. We introduce DAES, a training-free framework for Decoupled Adaptive Evidence Selection. DAES uses reflection-based cues to disentangle visual-text and ordinary visual evidence at the patch level, and estimates their query-conditioned relevance through Value-space similarity between query tokens and visual patches in an MLLM. To make the two evidence types directly comparable, we calibrate their relevance against type-specific null distributions. Under a fixed token budget, DAES adaptively selects evidence types and timestamps according to calibrated relevance scores. This evidence-level selection enables broader temporal coverage under fixed token budget while preserving question-relevant information. For efficient selection, we use a lightweight MLLM as the evidence selector while reserving a larger MLLM for final reasoning. Experiments on subtitle-rendered Video-MME-v2 and LongVideoBench demonstrate consistent improvements in video understanding performance under the same token budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.