acceptodds
Under review as a conference paper at ICLR 2027

StreaMem: Can MLLMs Understand Real-World Streams with Long-Term Memory?

Abstract

Deployment of multimodal large language models (MLLMs) demands systems that process continuous multimodal streaming inputs at a fixed rate, respond instantly, and retain dynamically evolving long-term memory, capabilities that current MLLMs still largely lack. To rigorously assess such capabilities, we build **StreaMem**, the first benchmark to jointly evaluate MLLMs' real-time processing and long-term memory under a real-streaming protocol, covering four scenarios and 12 subtasks with 210 videos of more than 30 minutes and over 3,000 open-ended QA pairs, together with four metrics. Evaluation of 7 existing models on StreaMem confirms that real-time processing and long-term memory remain unsolved under real-streaming conditions. We then propose **SOAM**, a **S**treaming **O**mni **A**gentic **M**emory that equips MLLMs with streaming understanding and external long-term memory without training. SOAM parses the incoming stream into semantically coherent events via event boundary detection, organizes them into a heterogeneous two-layer memory graph with biometric entity anchoring, routes each question through an LLM Router that decides whether memory retrieval is needed, and retrieves evidence along complementary semantic, entity co-occurrence, and temporal-anchor paths. SOAM achieves state-of-the-art performance among open-source models and methods of comparable scale, with substantial improvements in long-term memory persistence. The code is anonymously available at https://anonymous.4open.science/r/SOAM-FB1324wrc/, and the dataset will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.