Keeping Pace with Video: Online Multimodal Memory Construction
Abstract
Long-form video assistants must make incoming observations searchable during ongoing use, before future questions are known. This requires online memory construction that keeps pace with the video stream under a fixed compute budget. When construction falls behind, unfinished work accumulates and access to recent evidence is delayed. We formulate online memory construction as a joint quality-stability problem: supporting useful retrieval while keeping construction workload bounded under sustained arrivals. Building on the insight that linguistic indexing does not require generating descriptions, we introduce **LoreMem** (**L**exicon-**O**rdinal **Re**trieval **Mem**ory), a caption-free multimodal memory. Using frozen encoders,**LoreMem** selects concepts from a predefined lexicon to index visual observations and dialogue, storing sparse concept identities and ranks. The index provides semantic cues for locating evidence, while the original observations retain the details needed for answering. On completed indexes across four video QA benchmarks, **LoreMem** achieves 52.0% average accuracy, exceeding Qwen3-VL-4B caption memory by 7.9 percentage points. Separately, workload evaluation using construction times directly measured on a Galaxy S26 Ultra shows no backlog accumulation over six hours of 30-second arrivals.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.