Repairable Uncertainty: Selective Temporal Reacquisition for Long-Video Understanding with MLLMs
Abstract
Multimodal large language models (MLLMs) can process only a limited number of frames from long videos due to their large context requirements. Existing methods typically rely on uniform sampling or a separate frame selector. Uniform sampling can miss brief but decisive events, while selector-based pipelines often rely on a weaker selector than the base MLLM, potentially discarding useful evidence before it reaches the answerer. We observe that under partial evidence, an MLLM can remain uncertain about the answer while its temporal attention is concentrated on a small set of query-relevant regions. We call this repairable uncertainty: the answer remains uncertain, but focused attention provides a signal for where the missing evidence may lie. We introduce Selective Temporal Reacquisition (STR), a lightweight test-time method that uses a frozen MLLM’s answer uncertainty and temporal relevance to decide whether and where to reacquire video evidence. Starting from a uniform grid of \(F\) frames, STR uses the MLLM’s temporal attention to construct a new sampling density, triggering a budget-matched second pass only when the answer is uncertain and attention is sufficiently focused. We further introduce STR-Refine, which gates repeated observations using predictive uncertainty and accumulates relevance across them, and STR-Lite, which reduces the visual-token cost of reacquisition. Across five benchmarks (LongVideoBench, VideoMME, MLVU, VideoVista, and Video-MME-v2) and four MLLMs (Qwen3-VL-8B, Qwen3-VL-2B, Qwen3-VL-32B, and InternVL3-8B), STR consistently improves long-video understanding and achieves state-of-the-art performance. These results establish STR as a test-time framework for selectively reacquiring missing evidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.