acceptodds
Under review as a conference paper at ICLR 2027

EvoVideoRAG: Navigate, Verify, and Evolve for Token-Efficient Video Understanding via On-Policy Video Distillation

Abstract

Retrieval-Augmented Generation (RAG) grounds model generation in external evidence. With the development of multimodal models and agentic techniques, RAG has expanded from one-shot retrieval to continual evidence search, observation, and verification across heterogeneous information sources. Video is a particularly challenging modality: query-relevant information is temporally sparse and distributed across multiple scales, while each observation incurs substantial visual-token and computational costs. Existing video agents are commonly designed for individual tasks, making it difficult for a single policy to handle tasks with different input and output requirements, such as video corpus moment retrieval (VCMR) and video question answering (VideoQA). Meanwhile, successful and failed interaction trajectories are often discarded after a task, whereas continually accumulating external skills increases maintenance and invocation costs at inference time. To address these challenges, we propose EvoVideoRAG, a unified, task-adaptive, self-evolving, and cost-aware RAG framework for video-language understanding. Specifically, (1) EvoVideoRAG formulates heterogeneous tasks, including VCMR and VideoQA, as task-conditioned sequential evidence acquisition, enabling a shared agent to plan, retrieve, and verify video evidence and to adaptively decide whether to continue searching, adjust the temporal range, increase observation fidelity, or submit a prediction; (2) we introduce On-Policy Video Self-Distillation (OPVD), which turns new trajectories generated by the current policy into self-supervision for its successor, whose renewed interaction with the environment forms an intergenerational “interaction–distillation–reinteraction” self-evolution cycle without requiring an additional teacher or an ever-growing long-term skill library at inference time; and (3) we employ dynamic visual budgeting and cost constraints to allocate visual computation according to task difficulty and evidence sufficiency, reducing redundant frames and visual-token consumption while preserving task performance. Experiments on NExT-QA, Video-MME, Charades-FIG, and DiDeMo-FIG demonstrate the effectiveness of EvoVideoRAG in unifying video retrieval, temporal localization, and question answering through task-adaptive evidence acquisition within a unified agentic framework.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.