acceptodds
Under review as a conference paper at ICLR 2027

MEMLENS: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

Abstract

Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capability: longcontext LVLMs and memory-augmented agents. However, no existing benchmark conducts a systematic comparison of the two on questions that genuinely require multimodal evidence. To close this gap, we introduce MEMLENS, a comprehensive benchmark for memory in multimodal multi-session conversations, comprising 789 questions across five memory abilities (information extraction, multi-session reasoning, temporal reasoning, knowledge update, and answer refusal) at four standard context lengths (32K–256K tokens) under a cross-modal token-counting scheme. On 634 image-essential and image-supportive questions, controlled image ablation drops two frontier LVLMs below 2% accuracy. Across 27 LVLMs and 7 memoryaugmented agents, the best 32K accuracies are 58.68% for LVLMs and 33.46% for memory agents. LVLMs degrade as conversations grow, whereas memory agents are comparatively length-stable but lose visual fidelity under storage-time compression. Multi-session reasoning caps most systems below 30%, and neither approach alone solves the task. Abstention results suggest memory-oriented post-training risks weakening correct refusal. These results motivate hybrid architectures that combine long-context attention with structured multimodal retrieval.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.