acceptodds
Under review as a conference paper at ICLR 2027

RISE: Recurrent Evidence Consolidation with Kernelized Selection for Efficient Long-Video Understanding

Abstract

Although recent advancements in video large language models (Video LLMs) have enabled remarkable vision-language reasoning, long-video understanding suffer from the scalability issue due to substantial overhead of large-scale visual tokens. Existing approaches typically identify informative tokens based on local density or static saliency, retaining redundant semantically similar tokens while failing to sufficiently cover global video semantics. Towards this end, we propose a novel approach named Recurrent Evidence Consolidation with Kernelized Selection (RISE) for efficient long-video understanding. The core of our RISE is to compress redundant visual tokens from the views of both temporal consistency and information utility for efficient reasoning. In particular, our RISE first calculates the attention-based consistency scores to quantify recurrent patterns among spatio-temporally neighboring tokens in a sliding-window manner. These scores would be utilized to guide the anchor-based aggregation of evidence-consistent tokens from neighboring frames. In addition, we measure the conditional evidence gain according to a determinantal point process (DPP) kernel and iteratively select the most informative candidates from the remaining tokens. In this way, we maximize the information utility of the selected tokens within the constraints of computational budgets. Extensive experiments on four video benchmarks demonstrate that RISE consistently outperforms state-of-the-art baselines across various video LLM backbones. The code is available at https://anonymous.4open.science/r/RISE_4_Video_Compression

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.