acceptodds
Under review as a conference paper at ICLR 2027

IMPRINT: Rethinking Retrieval for Long-Term Conversational Agent Memory

Abstract

Long-term memory systems for conversational agents differ in how they organize memory but share one retrieval primitive: dense embeddings ranked by cosine similarity. We argue that this primitive, not the memory structure, is the bottleneck, because personal recall hinges on rare user-specific identifiers whose exact identity mean pooling averages away. We reframe recall as *imprint-following*, where each rare token is a trace pointing back to its source session, and instantiate it in **IMPRINT**, whose single *Token Profile Table* serves as an IDF-weighted inverted index and drives an online consolidation pool, with no retrieval training, no LLM calls at indexing, and no GPU at query time. We prove that rarity-weighted exact matching keeps a separation margin independent of the retrieval-unit length, whereas the dense margin decays as ; the same result explains why session-level sparse retrieval was dismissed by prior work. On LoCoMo (27 sessions) and LongMemEval (500 sessions), IMPRINT beats the strongest prior memory system by up to +15.8 Recall@3 and +3.9 QA-F1, outperforms dense retrieval under an identical pipeline, and produces the best answers on LoCoMo.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.