Cog-Fovea: Conditional Visual Acquisition and Answer Selection for Video Reasoning
Abstract
Sparse frame sampling can omit brief events essential for video reasoning. Extending textual reasoning may improve the interpretation of observed frames, but does not acquire missing observations. We introduce Cog-Fovea, an inference framework with frozen weights that separates two decisions: when to inspect additional video frames and when to replace an existing answer. An initial global pass at high resolution produces an answer and a confidence score. Queries with low confidence trigger temporal grounding and a hybrid view at lower resolution combining local samples with global anchors. Arbitration based on confidence selects between the original and hybrid answers, retaining the initial answer as a candidate rather than automatically accepting the new prediction. All stages share a frozen backbone, with a cap on visual tokens per pass and total computation accounting for routed calls. A conditional analysis of event coverage characterizes when localized sampling can capture brief events. Experiments on five benchmarks examine accuracy, inference cost, and answer selection behavior. On Video-MMMU, Cog-Fovea achieves 60.2% accuracy, compared with 57.0% using only the global pass. Inference without explanations averages 6.3 seconds per query. Approximately 35% of queries receive additional inspection. Arbitration provides a numerical gain of 0.6 percentage points over accepting every available hybrid answer, but reduces accuracy on a diagnostic subset with narrow temporal windows. These findings support conditional visual acquisition as a strategy for allocating computation at test time and show why acquiring more evidence and accepting a revised answer must be evaluated separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.