acceptodds
Under review as a conference paper at ICLR 2027

GraphCard: A Simple Graph Memory Plug-in for Long-Term Conversational Memory Systems

Abstract

Conversational memory in deployed systems is predominantly vector-based: memories are stored as extracted facts or raw dialogue chunks, and a query is answered through fast per-entry similarity search. Production favors this design for its speed, scalability, and strong overall accuracy. However, multi-hop and temporal questions depend on relations and timelines that live between entries, and per-entry retrieval cannot see them. A graph memory captures exactly this missing layer, linking entities through facts. The two designs are natural complements: vector retrieval buys coverage, graph structure buys relations. Their combination, however, remains underexplored: existing graph designs replace the vector stack rather than augment it, and we find they score far below the strongest vector-based systems they aim to improve. In this paper, we take the strengths of both by augmenting rather than replacing: the existing vector memory (the host) stays untouched, and GraphCard, a graph plug-in, runs beside it. At memory ingestion time, a plug-in branch runs parallel to the host's own pipeline, organizing the raw conversation stream into a fact graph through entity deduplication and temporal fact management. At search time, a single query embedding and a bounded graph search, with no LLM call, quickly condense the relevant subgraph into one  150-token graph card prepended on top of the host's retrieved memories. Nothing else in the host memory system changes, yet the missing relational layer is now in the context. Extensive experiments across two memory benchmarks (LoCoMo, LongMemEval), two model groups (GPT, Claude), and four strong vector-based baseline hosts (RAG, Nemori, SimpleMem, Mem0) show that GraphCard improves every host configuration, by up to +12.3 points, with gains concentrated on multi-hop and temporal questions. GraphCard also yields a better accuracy–token trade-off than widening the host's own retrieval: against Mem0, GraphCard is  1.5 points more accurate at a matched memory budget; it matches the same host's accuracy with 79% fewer additional tokens. GraphCard is fast as well, with an LLM-free search that completes in 0.61 s at p50, in the same sub-second range as the vector hosts themselves.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.