DSV-Mem: Evaluating multimodal memory in professional workflows for MLLM agents
Abstract
Conversational MLLM agents are increasingly expected to assist in professional workflows: from AI research and engineering design to product management and business operations. Despite this growing need, this capability has not been well explored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios that typically feature photographic natural images, isolated, static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, as well as compositional queries that require reconciling many artifact versions in memory while tracking state precisely. To address these challenges, we introduce DSV-Mem, a carefully curated benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises a diverse suite of expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). Moreover, a Hartley-inspired criterion is introduced to favor questions with broader visual-evidence inspection demands. To support further study, we introduce a generation harness that produces high-quality, domain-specific evaluation suites by decoupling state-transition synthesis from conversation filling. Rigorous evaluation over 27 configurations spanning frontier, open-weight models and popular memory management methods reveals that even the strongest baseline scores below 45% on DSV-Mem. Comprehensive analysis surfaces several important findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR and arithmetic are not the primary bottlenecks; 2) models frequently fail to verify user premises against prior state updates before answering; 3) popular interventions such as increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.