ForeSeek: Sizing and Front-Loading Multimodal Evidence for Long-Form Video Understanding
Abstract
Long-form video understanding is a core challenge in multimodal intelligence, where decisive evidence often rests on a few seconds within an hour-long video. Current retrieval-based pipelines typically caption the entire video in advance us- ing an auxiliary vision-language model, while agentic inspection systems gather observations through iterative coarse-to-fine reasoning. Although effective, in- dexed pipelines incur heavy duration-dependent captioning costs, while iterative agents defer evidence to later turns, by which point the model has already commit- ted to an answer it seldom revises. To overcome these limitations, we introduce ForeSeek, a retrieval-based framework that adaptively sizes and front-loads multi- modal evidence before answer commitment. ForeSeek first deploys the Channel- Aligned Evidence Assembler to align grounding frames, sparse visual descrip- tions, and verbatim speech on a shared timeline via a lightweight retriever. Con- ditioned on this timeline, the Pre-Answer Evidence Allocator reads model confi- dence using a fast few-token probe, front-loading a tailored package on the initial pass. Consequently, evidence volume adapts to query demand rather than video duration, confining subsequent interaction to bounded local inspection. Exten- sive experiments on multiple benchmarks demonstrate the effectiveness of our method, attributable to front-loading adaptively sized evidence packages through pre-answer confidence probing and channel-aligned multimodal assembly
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.