From Error Propagation to Recovery: Principles for Persistent Memory in Multi-Agent Systems
Abstract
Error propagation undermines reliable decision-making in large language model-based multi-agent systems (LLM-MAS). Existing work uses persistent memory to mitigate this problem, but **lacks a theoretical characterization of how memory enables recovery**. This leaves memory design without theoretical guidance on **what to remember** and **when to intervene**, potentially leading to unnecessary memory use and increased **token and storage costs**. To address this gap, we connect error propagation and recovery analysis, establishing the necessity of persistent memory in LLM-MAS under information loss and deriving principles for memory design. Specifically, we derive a closed-loop error propagation law coupling agent communication and environment feedback. We further analyze error recovery under information loss and show that persistent memory reduces the optimal error from random guessing to zero. Guided by this theoretical analysis, we develop , an in-task memory for efficient error recovery. Experiments across five benchmarks and three MAS frameworks demonstrate performance gains over in-task and cross-task memory baselines, with efficiency analyses showing reduced token usage and storage overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.