acceptodds
Under review as a conference paper at ICLR 2027

Efficient Multimodal Memory Construction for Long-Form Video Retrieval

Abstract

Long-form video question answering requires retaining sparse evidence across extended temporal contexts. Retrieval-based systems make this setting tractable, but they can only retrieve evidence that was preserved when the video memory was written. Under a lightweight memory builder, compressing each clip into one query-agnostic caption is a difficult target: a generic narrative can omit fine-grained cues that later become relevant. We therefore propose **MVP-Mem**, a retrieval-oriented framework that represents each clip as short phrases extracted under complementary semantic views. The resulting phrases form compact, independently retrievable evidence anchors, accessed through a simple phrase-wise maximum scoring rule. Our primary comparison uses Qwen3-VL-4B for both caption and phrase memory construction, with the same downstream answer generator. Across four long-form video QA benchmarks, **MVP-Mem** improves over this caption baseline on every benchmark, raising average accuracy from 44.5% to 54.0% (+9.5 percentage points). On EgoLifeQA, it also nearly halves memory construction time relative to caption generation with the same Qwen3-VL-4B builder. These results support multi-view phrase memory as an effective construction target for lightweight VLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.