acceptodds
Under review as a conference paper at ICLR 2027

CSMem: Cognitive State Memory for Multimodal Long-Term Dialogue Assistants

Abstract

Multimodal long-term dialogue memory asks assistants to answer later questions from past image-text interactions. Existing multimodal memory methods preserve useful content and evidence through textualized records, raw multimodal memories, or graph-structured representations. However, these methods primarily preserve what appeared, rather than explicitly modeling the interaction state that gives those records meaning in context, including how references are resolved, alternatives are compared, decisions evolve, and claims are supported. To address this limitation, we introduce Cognitive State Memory (CSMem), a training-free memory framework. Inspired by situation-model accounts of comprehension, CSMem represents the structured state established within each coherent interaction rather than storing only its surface image-text records. It connects this state with retrieval cues and raw multimodal evidence to locate relevant interactions and generate grounded answers. Experiments on Mem-Gallery and LoCoMo show that CSMem outperforms textualized and raw multimodal methods, achieving up to an 11.3% relative improvement in exact match. Our code will be made publicly available soon.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.