From Similar to Useful: Failure-Attributed Memory Selection for LLM Agents
Abstract
LLM-based agents increasingly rely on memory of past experience to adapt without updating their parameters. Existing work concentrates on the memory bank itself, addressing how agent trajectories should be processed and how the bank should be maintained. Less attention has been paid to a different question: why memories that appear relevant often fail to help, and what makes a memory useful to a given agent on a given task. We study this question in a controlled setting where the memory bank holds only verified successful trajectories, leaving selection as the only variable. For each task the agent fails without memory, we inject each retrieved candidate separately and examine how recovery relates to the agent's failed trajectory. Across three datasets covering code generation, embodied interaction and text-to-SQL, even among highly similar tasks, how the agent failed carries information about which memory will help, and memories that demonstrate the behavior a failure lacks are more likely to recover it. Building on this, we propose FAMA (Failure-Attributed Memory Allocation), an LLM-based selector that reads the agent's failed rollout, identifies what it lacks, and chooses the memory that addresses that gap. On the three datasets, under both single-agent and multi-agent architectures, fine-tuned FAMA recovers 8 to 11 more points of failed tasks than top-1 similarity retrieval. It also outperforms a cross-encoder reranker, a selector that does not observe the failure, and Reflexion under the same retry budget and partly transfers to an executor it was not trained on.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.