VideoLedger: Index-and-Revisit Memory for Long-Horizon Video Stream Understanding
Abstract
Large multimodal models (LMMs) remain limited when reasoning over continuously growing visual histories spanning hours, weeks, or months. Reprocessing the accumulated video for every new question becomes increasingly impractical, while relying on compact memory alone risks losing details that may become decisive for future questions. We introduce **VideoLedger**, a training-free framework that follows an *index-and-revisit* paradigm for long-horizon streaming video understanding. As each clip arrives, VideoLedger appends timestamped visual narratives and audio transcripts to chronological Markdown timelines, forming a question-agnostic index linked to the source video. At question time, an off-the-shelf LMM navigates the index using only standard file operations—lookup and scoped reading—and selectively revisits timestamp-linked source clips when the textual records are insufficient. This design separates efficient long-range evidence navigation from fine-grained visual verification. Stored as persistent text outside the active reasoning context, the index is naturally harness-friendly and remains accessible after context compaction, without a learned retriever or graph traversal. Notably, this simple navigation interface proves highly effective: across ten hour-, week-, and month-scale benchmarks, VideoLedger achieves the best results among compared methods in ultra-long, robot-centric, and general long-video settings, with consistent gains across Qwen, GPT, and Gemini backbones. Further analyses show complementary benefits from textual and revisited visual evidence, accurate temporal grounding, and a favorable accuracy–latency trade-off. Our code and evaluation logs will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.