acceptodds
Under review as a conference paper at ICLR 2027

Memory Use in Language-Model Agents: Action Recommendations and Status Claims

Abstract

In applications such as document delivery and task coordination, language-model agents rely on past interactions to determine what to do next and which actions have already been completed. Retrieval performance alone, however, does not show whether agents use the available evidence appropriately. We study the gap between acquiring memory evidence and using it to recommend actions, distinguishing current-state evidence completeness, task satisfaction, and response appropriateness given the visible evidence. We combine longitudinal retrieval diagnostics across six memory-system configurations with controlled offline evaluations of two language models in scenarios involving version updates, authorization changes, and execution records. Our experiments vary retrieval budgets, repair missing evidence, and compare history access and prompts. Action recommendations can remain incorrect even when retrieval supplies all required evidence. Targeted evidence repair resolves some failures, but retrieving more records does not consistently help. Access to full history improves task satisfaction for both models even when the same task content is available without history, so the observed gains cannot be explained solely by access to that content. Further analysis distinguishes task satisfaction from response appropriateness: a reasonable clarification request may leave task requirements unmet, whereas a recommendation that satisfies them may accompany unsupported claims of completed actions. Prompt changes do not yield consistent improvements across models. These results show why evaluating memory-supported agents requires assessing both the adequacy of retrieved evidence and how agents use it in recommendations and status claims.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.