acceptodds
Under review as a conference paper at ICLR 2027

FocusOmni: Coarse-to-Fine Temporal Evidence Grounding for Omni-Modal Video Reasoning

Abstract

Large vision-language models (VLMs) have made significant progress in video reasoning by integrating visual, audio, and textual modalities. However, when long videos contain substantial irrelevant content, existing models that directly reason over uniformly sampled frames often suffer from temporal distraction, as they lack an explicit mechanism to identify which temporal evidence should participate in reasoning. To address this limitation, we propose FocusOmni, a coarse-to-fine temporal evidence grounding framework that enables omni-modal models to identify question-relevant temporal regions before performing multimodal reasoning, which follows a coarse-to-fine grounding strategy. It first constructs a compact global temporal overview by arranging sampled frames into a grid representation, from which the model identifies the temporal region that contains question-relevant evidence. Based on the grounded temporal region, FocusOmni concentrates visual and auditory evidence within the relevant interval, enabling more effective omni-modal reasoning through explicit temporal evidence grounding. We further introduce ConcatBench, a benchmark that constructs temporally distracted videos by combining question-relevant clips with irrelevant distractor clips, enabling evaluation of segment-level temporal evidence grounding under controlled interference. Extensive experiments on ConcatBench and additional video benchmarks demonstrate that FocusOmni improves both temporal grounding and downstream reasoning performance while requiring only a compact global overview and localized evidence frames.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.