acceptodds
Under review as a conference paper at ICLR 2027

Less Context, More Accuracy: Structured Long-Term Memory Beyond Equal-Budget RAG

Abstract

Long-term memory for LLM agents is often justified against replaying the entire interaction history, which is expensive and can become less accurate as distractors accumulate. Yet outperforming full context does not establish that structured memory adds value beyond ordinary retrieval: a token-matched RAG baseline may recover the same gain. Existing evaluations also change the reader, judge, or context budget together with the memory system, making gains difficult to attribute and reproduce. We present Engram, an open-source, dual-process memory engine built on a bi-temporal data model. A fast write path appends lossless episodes without an LLM on the critical path; an asynchronous consolidation path extracts atomic (subject, predicate, object) facts, builds a bi-temporal knowledge graph, and resolves contradictions without an LLM call per fact—invalidating, never deleting, so every fact retains provenance and a supersession chain. A hybrid read path fuses dense, lexical, graph, and recency/salience signals, applies a point-in-time (“as-of”) temporal filter, and assembles a compact, provenance-tagged context. We evaluate frozen configurations on all 500 questions of LongMemEval-S with the official category-specific judge prompts. Streamlined Engram reaches 78.8%, significantly outperforming an equal-budget dense+BM25 RAG control at 72.4% (+6.4 points; McNemar p=0.001) and full context at 61.6% (+17.2 points), while using 6.7k rather than 103.5k mean context tokens (15.5x compression). The 86.0% annotated-evidence oracle is +7.2 points higher. The predefined 440-question held-out partition preserves the margins. Five single-variable interventions show that gains are category-specific and that an evidence planner and hierarchy blocks reduce aggregate accuracy; repeated generations and a 2x2 reader-judge cross test robustness. A separate streaming experiment on all 500 LongMemEval-M questions measures the shared episode index without conflating retrieval recall with QA accuracy. Every reported value is generated from fingerprinted, append-only context, answer, and judgment artifacts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.