acceptodds
Under review as a conference paper at ICLR 2027

CERES: RECONCILING DURABLE EVIDENCE AND EVOLVING ACTIVE STATE FOR LONG-LIVED LLM AGENTS

Abstract

Large language models (LLMs) are increasingly used to build long-lived agents, which rely on external memory to support their long-term operation. However, the core challenge of memory goes beyond retrieval: a system must preserve raw interaction evidence before memory compression. It must also revise existing records and preserve the basis for each revision as facts change, reconcile conflicting records, and prevent outdated conclusions from continuing to dominate agent behavior. Many memory methods emphasize retrieval, while preserving information during writing, updating facts over time, and resolving conflicts remain challenging. Omitted evidence may be irrecoverable downstream, contradictions may persist, and higher-order conclusions synthesized from interaction histories may become outdated without active maintenance. We introduce CERES (Continual Evidence Reconciliation and Evolving State), which unifies these operations as an evidence-grounded memory lifecycle. CERES comprises two core components. Engram first preserves raw interaction evidence, then extracts recall-worthy information into structured records and integrates these records into structured memory pages with source links. Cortex maintains a capacity-limited set of higher-order representations grounded in these pages. It admits, refreshes, or retires entries based on changes in their supporting evidence, recorded usage and feedback, making them available as default-active context. Across end-to-end comparisons on Memora and an auxiliary evaluation on LongMemEval-S, CERES demonstrates more reliable long-term memory performance. Write and recall experiments further show that Engram consolidates cross-session evidence into reliable memory that is traceable and aligned with the current state. Under identical memory content, comparing answering with and without Cortex guidance shows that it improves average answer quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.