acceptodds
Under review as a conference paper at ICLR 2027

AttestMem: Turning Long Conversations into Traceable, Revisable Memory

Abstract

Long-horizon language agents require persistent memory. Existing systems reduced dialogue histories to collections of independently embedded chunks: outdated or contradictory information can coexist at retrieval time, and semantic similarity alone decides which memories enter the answering context. We introduce AttestMem, a training-free memory framework that represents conversations as a revisable collection of attested statements and compiles a bounded answering context from recalled candidates. Each statement carries its extraction provenance, utterance time, validity interval, and ingestion order; when deterministic source matching succeeds, the row additionally carries turn-level references and its source-match grade, and ambiguous or unsupported cases are explicitly marked rather than guessed. An append-only event stream incorporates updates so that they supersede rather than overwrite earlier values, while rejected writes and historical versions remain queryable. At inference time, candidates are nominated through complementary routes and compiled under a character budget that prioritizes assertions and recovered source quotations while trimming transcript text. With GPT-4o-mini for memory construction and answering, AttestMem reaches 89.03% accuracy on LoCoMo and 86.00% on LongMemEval-S; with GPT-4.1-mini, 93.44% and 89.10%, using 3.6–3.9k prompt tokens per question. These results match or exceed the strongest reported baseline in each backbone group on both benchmarks. Paired ablations show the largest accuracy decrease when the ledger is removed, and removing the context limit lowers LoCoMo accuracy by 14.25 points while tripling the prompt.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.