acceptodds
Under review as a conference paper at ICLR 2027

Beyond Positive Reinforcement: A Hierarchical Multi-Perspective Memory Framework for Multi-Agent Systems

Abstract

Recent advancements in Multi-Agent Systems (MAS) have witnessed significant progress in equipping agents with episodic memory, typically designed to store and retrieve successful collaborative trajectories. However, a critical design consideration remains insufficiently addressed in existing approaches: they predominantly rely on positive reinforcement, with insufficient utilization of the pedagogical value provided by failure experiences. Ignoring these error signals leads to a positive bias, making agents prone to local optima and less robust against cascading failures. To bridge this gap, we propose a Multi-perspective Memory (M-Memory) framework that draws inspiration from three empirically-grounded memory organization strategies. Concretely, M-Memory integrates both successful and failed memory records and structures them at two levels. At the high level, the framework synthesizes three types of insights: Shortcuts from successful tasks, Pitfalls from failures, and Cross-domain Consensus from diverse experiences. At the low level, it enriches key interaction steps with trajectories from both failures and varied successes to boost robustness and generalization. For a new task, M-Memory bidirectionally retrieves high-level insights and fine-grained trajectories to guide execution. By integrating new outcomes, this hierarchical feedback loop drives incremental MAS evolution. The superiority of our approach is demonstrated through extensive experiments across four benchmarks, two LLM architectures, and three prominent MAS.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.