acceptodds
Under review as a conference paper at ICLR 2027

IR³: Index Once, Retrieve, Restore, and Reduce for Multi-Query Long-Video Understanding

Abstract

Popular long videos often receive many questions from different users, yet conventional systems repeatedly perform query-independent visual processing for every question, which leads to substantial redundant computation and inefficient online processing. Existing efficiency approaches mainly follow two directions. Query-conditioned keyframe retrieval improves evidence relevance, but typically introduces additional online scoring and selection latency and may fragment temporal context, omitting event development or procedural order. Visual-token compression, in contrast, reduces downstream inference cost by removing redundant visual tokens while largely preserving task performance, but usually operates only on already selected evidence. Moreover, learning fine-grained token-retention decisions from answer-level supervision creates a difficult credit-assignment problem. We introduce (**I**ndex Once, **R**etrieve, **R**estore, and **R**educe) for efficient multi-query long-video understanding. constructs a reusable, query-independent visual index once per video and uses a hybrid global–local retriever to efficiently locate sparse question-relevant anchors. It then restores their contiguous temporal context and compresses the completed evidence to exactly match the visual-token budget of direct sparse inference. To address diffuse credit assignment, introduces semantic-block budget allocation, enabling more structured and stable token-budget distribution across semantically coherent regions. A causal memory further guides subsequent allocations using evidence retained from earlier visual slices. Across 6,196 questions from three long-video benchmarks, 's retrieval-only path maintains nearly identical accuracy to WFS32 while reducing steady-state online latency from 8.450 to 0.510 seconds. On videos of at least 30 minutes, provides accuracy gains while preserving the same final visual-token budget. Matched-training ablations further show that semantic-block actions provide more stable optimization than token-wise actions under answer-level supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.