acceptodds
Under review as a conference paper at ICLR 2027

VideoMem-Bench: Benchmarking Query-Agnostic Memory in Video Understanding

Abstract

While recent Video Large Language Models (Video-LLMs) show strong performance in processing long context windows, it remains unclear whether they genuinely memorize information or merely retrieve it from a large context cache. Existing benchmarks often overestimate actual memory capabilities because they rely on query-guided joint encoding, allowing models to retrieve video frames conditioned on the question. Furthermore, these benchmarks lack a structured cognitive hierarchy. To address these limitations, we introduce VideoMem-Bench, a benchmark designed to systematically evaluate genuine memory in the field of long video understanding. We establish a four-level hierarchical framework inspired by cognitive neuroscience, ranging from basic semantic perception to predictive causal simulation. Crucially, we propose a query-agnostic, decoupled evaluation paradigm. By enforcing an information bottleneck, this paradigm requires models to compress continuous video streams into a capacity-constrained memory state without prior knowledge of the queries. Models must then answer subsequent questions relying solely on this compressed memory representation. Comprising 3,433 videos and 5,360 annotated QA pairs, VideoMem-Bench prevents retrieval shortcuts and provides a rigorous standard for assessing autonomous long-term memory in dynamic environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.