Beware of Memory-Induced Overconfidence: Commitment–Competence Calibration for Multi-Agent Systems
Abstract
Large Language Model-based Multi-Agent Systems have demonstrated impressive capabilities in solving complex tasks through role specialization, memory sharing, and collaborative reasoning. To further enhance their long-term performance, recent studies have introduced self-evolution mechanisms, where agents summarize reflections from historical interactions and reuse them in subsequent tasks. However, we observe that self-evolution does not always lead to consistent improvement: in some cases, apparently relevant memories may provide little benefit for the current task while repeatedly steering agents toward the same ineffective strategy. The key insight is that memory injection has two distinct effects: memory commitment, which characterizes how strongly memory influences behavior by reducing uncertainty about how to act, and memory competence, which estimates how reliably the selected memories support successful execution. These effects can be misaligned: memory may impose a clear and consistent behavioral direction while exhibiting low historical execution reliability. We identify this failure mode as memory-induced overconfidence. To mitigate this challenge, we propose C2Mem, a memory calibration framework that dynamically attenuates apparently relevant memories that exhibit low execution utility. Experiments across multiple benchmarks show that C2Mem reduces the propagation of non-beneficial memory guidance, mitigates memory-induced overconfidence, and improves the stability and reliability of multi-agent systems. Our code is available at https://anonymous.4open.science/r/c2mem-27F3.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.