acceptodds
Under review as a conference paper at ICLR 2027

Remember, Then Decide: How Stored State and Readers Shape Premise Correction in Agent Memory

Abstract

Agent memory must correct questions that assume outdated facts while answering valid questions and recalling current and historical information. We introduce LEDGER, an evidence-preserving memory that retains raw dialogue and attribute histories, and evaluate how stored state and answering procedures affect these goals across eight systems and three benchmarks. Holding LEDGER evidence fixed, replacing direct answering with its query-dependent reader raises GPT-4.1 premise resistance from 0.020 to 0.725, although effects vary across stores and backbones. Matched queries over 117 histories nevertheless reveal that higher stale-premise accuracy can accompany lower valid-premise accuracy. On 131 held-out histories, oracle replacement of current values improves both accuracies on both tested backbones. Further replacing transition records with canonical summaries improves valid-premise accuracy on GPT-4.1 but reduces it on GPT-4o-mini. These findings show that correction depends on both the supplied state representation and the reader interpreting it. Preserving evidence and scoring well on separate benchmarks do not establish reliable selective correction. We therefore advocate evaluating stale- and valid-premise accuracy alongside current and historical recall, distinguishing premise rejection from answer correctness, and separating memory-construction effects from reader effects.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.