VideoSEAM: Stitching Coherent Video-Level Memory for Long-Form Video Understanding
Abstract
Long-form video understanding requires reasoning over temporally distributed entities, events, and interactions. Existing training-free methods broadly follow either query-conditioned approaches, often relying on sparse frame search that may miss critical moments, or query-independent approaches, where independently processed local windows can lead to inconsistent identities and fragmented events. We introduce VideoSEAM, a training-free framework for query-independent sequential video memorization. Rather than constructing local records independently and recovering their relations downstream, VideoSEAM maintains a coherent entity-event state during sequential processing and directly produces structured records with consistent cross-window references. A dual-horizon memory separates a bounded multimodal transition memory for continued understanding from a persistent episode memory that preserves structured history for future queries. The resulting memory supports a memory-grounded reasoning agent for retrieval, bounded video inspection, and answer generation. Experiments on MM-Lifelong, EgoLifeQA, and LVBench demonstrate strong performance against existing training-free long-video methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.