RLVU: Recursive Long Video Understanding With Compact Models
Abstract
Using frontier models such as GPT-5 for video question answering can be prohibitively expensive at scale, making cost an important consideration alongside accuracy in applications such as surveillance, online videos, and recorded meetings. This motivates the use of compact reasoning backbones such as GPT-5 mini and Qwen3.5-27B. However, these models often struggle to locate and connect relevant evidence scattered throughout long videos. To address this challenge, we propose RLVU, which combines reusable memory and then recursively reasons over this memory to locate and integrate relevant evidence. Specifically, RLVU constructs video memory through top-down temporal refinement and bottom-up entity consolidation: event level context is propagated to shorter clips, while observations of recurring entities are merged into global records. Given a question, RLVU then performs recursive reasoning over the constructed memory through top-down question decomposition and bottom-up evidence integration. Compared with direct inference using the compact model like GPT-5 mini, the base RLVU configuration improves accuracy by 11.2%, 12.4%, and 4.3% on LongVideoBench, LVBench, and Video-MME, respectively. RLVU also consistently improves accuracy across all three benchmarks when using Qwen3.5-27B as the reasoning backbone, compared with direct inference using the same model. These results demonstrate that structured video memory and recursive reasoning enable compact reasoning models to more effectively locate, verify, and integrate evidence across long videos. The code is at https://anonymous.4open.science/r/RLVU-4B58.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.