acceptodds
Under review as a conference paper at ICLR 2027

Remembering Is Not Recovering: Benchmarking Agent Memory for Task Recovery

Abstract

In real-world deployments, *large language model* (LLM) agents may encounter unexpected interruptions, such as network outages or machine failures, making reliable task recovery essential. However, existing agent memory benchmarks focus on using past experience to improve future tasks, rather than recovering the state of an interrupted task and resuming it without repeating completed work or skipping necessary steps. To address this gap, we introduce *RecoverMemBench*, which evaluates recovery rates and additional costs relative to uninterrupted runs under two settings: *task resumption*, which preserves progress and tests task state reconstruction, and *task rollback*, which restores an earlier environment state and tests whether memory identifies lost progress while retaining reusable information. Across 11 existing memory systems, average recovery rates remain below 50% in both settings, with over 150% additional steps in the rollback setting, highlighting substantial reliability and efficiency gaps. Our analysis reveals two key limitations: (1) semantic retrieval struggles to identify the information needed for recovery among similar execution records of the same task; and (2) inadequate memory state management causes agents to mistake work undone by rollback for completed work. Motivated by these, we develop a memory system that improves recovery rates by over 35% in both settings while reducing average additional cost by more than half. Overall, our work provides a foundation for evaluating and improving agent memory for reliable and efficient task recovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.