CSMem: Cognitive State Memory for Multimodal Long-Term Dialogue Assistants
Abstract
Multimodal long-term dialogue memory asks assistants to answer later questions from past image-text interactions. Existing multimodal memory methods preserve useful content and evidence through textualized records, raw multimodal memories, or graph-structured representations. However, these methods primarily preserve what appeared, rather than explicitly modeling the interaction state that gives those records meaning in context, including how references are resolved, alternatives are compared, decisions evolve, and claims are supported. To address this limitation, we introduce Cognitive State Memory (CSMem), a training-free memory framework. Inspired by situation-model accounts of comprehension, CSMem represents the structured state established within each coherent interaction rather than storing only its surface image-text records. It connects this state with retrieval cues and raw multimodal evidence to locate relevant interactions and generate grounded answers. Experiments on Mem-Gallery and LoCoMo show that CSMem outperforms textualized and raw multimodal methods, achieving up to an 11.3% relative improvement in exact match. Our code will be made publicly available soon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.