CoReMem: Context-Aware Memory Construction and Query-Aware Retrieval for Multimodal Conversations
Abstract
Long-horizon multimodal conversational agents must recover relevant visual evidence from extended interaction histories. Existing memory systems typically convert images to text, potentially omitting visual details, or retrieve complete images using global embeddings, which can underrepresent small but relevant regions. Their retrieval procedures also commonly rely on a fixed top- budget, which may omit evidence needed for multi-hop reasoning. We present CoReMem, a multimodal memory framework for context-aware construction and query-aware retrieval. During construction, a vision-language model uses only the dialogue observed so far to select potentially useful regions, SAM3 produces the corresponding masks, and validated regions become independently indexed subimage documents linked to their parent turns. During retrieval, hybrid text–image ranking first retrieves an initial ranked set of documents drawn from at most dialogue turns. A VLM checks whether this initial context is sufficient for the user query. If not, CoReMem constructs visual queries conditioned on the query and images from the retrieved turns. It marks retrieved subimages as candidate query regions and uses DINOv3-based image-to-image retrieval to expand the context. The sufficiency check is repeated after each expansion to determine whether further evidence is needed. On MemEye with Qwen3.6-27B, CoReMem achieves LLM-as-a-Judge score of , outperforming the strongest evaluated baseline by . It also improves performance by – on settings with longer dialogue and more topic transitions, demonstrating the value of context-aware memory construction and query-conditioned visual expansion for long-horizon multimodal conversation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.