MemVista: Benchmarking Long-Term Memory for Multimodal Agents in Evolving Projects
Abstract
Multimodal agents working on long-running projects must move beyond traditional memory compression and retrieval methods to dynamically track evolving project states. We present MemVista, a benchmark designed to evaluate long-term multimodal memory in project-oriented interactions spanning multiple interdependent sessions. Our automated pipeline utilizes project-specific state schemas and sequences of state operations to guide data generation and establish ground truth for memory evaluation. Observable state facts are embedded in images as fine-grained attributes and relationships, with state-driven edits applied to maintain visual continuity over time. LLM-simulated users progressively reveal state facts through multi-session interactions with a tool-augmented agent, generating conversation histories that interleave user text and images with multimodal tool outputs. We define four task families to assess recall of currently valid facts, reconstruction of changes and their stated causes, abstention when a fact is not yet established or is no longer valid, and multi-hop reasoning across interrelated facts. We systematically compare visual memory representations and context-management strategies across long-context, retrieval-based, and agentic memory systems. MemVista provides a controlled testbed for investigating how agents maintain, update, and reason over multimodal information throughout evolving projects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.