GraphCard: A Simple Graph Memory Plug-in for Long-Term Conversational Memory Systems
Abstract
Conversational memory in deployed systems is predominantly vector-based: memories are stored as extracted facts or raw dialogue chunks, and a query is answered through fast per-entry similarity search. Production favors this design for its speed, scalability, and strong overall accuracy. However, multi-hop and temporal questions depend on relations and timelines that live between entries, and per-entry retrieval cannot see them. A graph memory captures exactly this missing layer, linking entities through facts. The two designs are natural complements: vector retrieval buys coverage, graph structure buys relations. Their combination, however, remains underexplored: existing graph designs replace the vector stack rather than augment it, and we find they score far below the strongest vector-based systems they aim to improve. In this paper, we take the strengths of both by augmenting rather than replacing: the existing vector memory (the host) stays untouched, and GraphCard, a graph plug-in, runs beside it. At memory ingestion time, a plug-in branch runs parallel to the host's own pipeline, organizing the raw conversation stream into a fact graph through entity deduplication and temporal fact management. At search time, a single query embedding and a bounded graph search, with no LLM call, quickly condense the relevant subgraph into one 150-token graph card prepended on top of the host's retrieved memories. Nothing else in the host memory system changes, yet the missing relational layer is now in the context. Extensive experiments across two memory benchmarks (LoCoMo, LongMemEval), two model groups (GPT, Claude), and four strong vector-based baseline hosts (RAG, Nemori, SimpleMem, Mem0) show that GraphCard improves every host configuration, by up to +12.3 points, with gains concentrated on multi-hop and temporal questions. GraphCard also yields a better accuracy–token trade-off than widening the host's own retrieval: against Mem0, GraphCard is 1.5 points more accurate at a matched memory budget; it matches the same host's accuracy with 79% fewer additional tokens. GraphCard is fast as well, with an LLM-free search that completes in 0.61 s at p50, in the same sub-second range as the vector hosts themselves.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.