acceptodds
Under review as a conference paper at ICLR 2027

Beyond Answering: Small Models Empower Large Models for Long-Video Understanding

Abstract

Recent advances in long-video reasoning enable models to actively search for visual evidence through tool use. However, these approaches commonly assign evidence acquisition and answer generation to the same model, coupling repeated visual exploration with that model’s computational cost and reasoning limitations. We find that small models can more readily locate and verify relevant evidence yet struggle to integrate it into a correct answer, revealing distinct capability requirements for evidence acquisition and reasoning. Building on this observation, we propose an evidence-driven collaboration framework in which a small model handles repeated, multi-step evidence search and verification, while a large model integrates the acquired evidence and produces the final answer in a single invocation. Crucially, we distinguish evidence search under sparse observations from direct evidence localization: when decisive details remain unobserved, the true evidence location may not be a search target supported by the available cues. We therefore supervise search priorities grounded in the current observations, teaching the small model where to inspect next, while dense local verification determines each candidate's actual evidential value. This decoupled framework supports flexible test-time scaling through configurable evidence acquisition and answer routing. Experiments demonstrate that small models can strengthen large-model reasoning through evidence acquisition, extending the collaborative value of small models beyond answering. Pairing a 4B evidence search and verification model with a frozen 27B model for final reasoning achieves 74.8% on VideoMME, 79.2% on MLVU, and 58.7% on LVBench. Code, data, and models will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.