Distributed Evidence Memory: Deferring Irreversible Decisions for Reliable Long-Term Memory
Abstract
Long-term memory is essential for language-model agents operating over interaction histories beyond the context window. Existing memory systems typically compress past experiences into summaries, facts, embeddings, or structured representations, and then use a centralized mechanism to decide which memories enter the current context. We argue that this central-gating paradigm can create an irreversible evidence bottleneck: relevant memories may be excluded before their original content is read, conflicting versions may be resolved too early, and retrieval may stop before the necessary evidence is complete.We introduce Distributed Evidence Memory (DEM), which defers these decisions to query time. DEM lets raw-memory shards report query-conditioned evidence through a shared workspace, preserves conflicting versions with a Conflict Directed Acyclic Graph (CDAG), and uses Good-Turing Coverage Estimation (GTC) to stop retrieval based on evidence saturation rather than fixed budgets or model confidence. Across FactCon, LoCoMo, and LongMemEval, DEM consistently outperforms strong retrieval and memory baselines. On FactCon, DEM improves SubEM from 41.5 to 64.0 with DeepSeek-v4-flash, with the largest gains appearing in multi-hop, conflict-heavy, and large-memory settings. These results support a broader principle for long-term memory design: preserve uncertainty during storage and defer irreversible decisions until sufficient query-time evidence is available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.