acceptodds
Under review as a conference paper at ICLR 2027

LongVidGraph: Progressive Graph Memory for Retrieval-Augmented Omni-Modal Long-Video Understanding

Abstract

Long-video understanding with omni-modal models remains challenging due to limited context budgets and long-range temporal dependencies. Existing methods rely on dense frame sampling or simple context accumulation, which fail to maintain coherent information over long durations. We propose LongVidGraph, a retrieval-augmented framework with a progressive graph memory. The model incrementally constructs a structured graph of entities, actions, and dialogues from audio-visual signals in long videos. This graph serves as a compact and temporally consistent memory. To enable the model to learn this incremental and structured representation, we build a large-scale clip-graph supervision pipeline that produces aligned video-graph training data at scale. We fine-tune Qwen2.5-Omni on this data for progressive graph construction. At inference time, we perform query-conditioned retrieval over the graph memory to obtain relevant structured evidence. The retrieved evidence is used to support downstream understanding tasks. Experiments on multiple long-video benchmarks show consistent improvements over strong open-source baselines, with more significant gains on long-duration videos. The results demonstrate the effectiveness of progressive graph memory and retrieval for long-video understanding with omni models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.