acceptodds
Under review as a conference paper at ICLR 2027

Read First, Look on Doubt: A Text-First Framework for Long-Video Question Answering

Abstract

Multimodal large language models have advanced long-video question answering chiefly by looking harder, through longer frame contexts, compressed visual tokens and multimodal memory agents, all treating the collapse of visual coverage as the problem to solve. We report two observations that point the other way. As videos lengthen, the language of the footage, its speech, narration and on-screen text, answers a growing share of what is asked on its own, and accuracy from the language track alone grows steadily stronger with duration while accuracy from visual input collapses. Visual information added on top of the text in turn loses its marginal value, which falls from points on the shortest clips to at the longest tier. In response, we propose ReadFirst, a text-first framework that inverts the standard allocation. It distils every channel of a video into a multi-scale textual memory offline and answers by reading at query time, escalating to a targeted visual inspection only when two derivations of the same answer disagree, thereby spending perception only where the text runs out. Extensive experiments demonstrate that ReadFirst achieves state-of-the-art performance, reaching an average accuracy of 74.1 across five long-video benchmarks while resolving over 90% of questions in text alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.