-Observe: A Difference-Aware, Delayed-Commitment, Dual-Route Observation Agent for Efficient Video Question Answering
Abstract
Video question answering still lacks a reliable way to identify the visual evidence that distinguishes competing answers under a limited observation budget. Relevance-based selection may capture the question topic yet omit the state, order, count, or relation needed to determine the answer. A separate bottleneck arises when textual intermediates omit visual details available in the original frames. To address these limitations, we present D-Observe, a training-free observation agent that uses the question and all answer options to select visual evidence before committing to an answer. It then passes the retained RGB frames directly to Vision-Language Models (VLMs) in a single call. Additionally, a label-free shared-observation route amortizes visual processing across multiple questions about the same video. Across five benchmarks spanning long and short videos (Video-MME, MLVU, EgoSchema, NExT-QA, and LVBench), D-Observe achieves 64.94% macro accuracy with an average of 10.21 frames sent per question. It improves upon uniform 32-frame inference with Qwen3.5-27B by 2.02 percentage points (95% CI [, ]) while reducing sent frames by 68.1%. By holding frame inputs fixed, we further isolate the effect of the textual bottleneck and show that direct visual answering consistently outperforms text-mediated answering.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.