Toward robust llm-based multi-agent systems in communication-limited environments
Abstract
The success of LLM-based Multi-Agent Systems (LLM-MAS) relies on reliable communication among agents. However, real world network fluctuations, execution timeouts, and API limits can cause unpredictable agent dropouts, removing agents and their associated communication links and thereby disrupting the communication topology. Existing studies have considered communication optimization and agent dropout resilience, but typically do not jointly address topology reconstruction according to the realized dropout pattern and the use of recovery history for subsequent adaptation. We investigate the impact of agent dropout on LLM-MAS and find that both spatial topology structure and temporal recovery dependence are critical for maintaining effective coordination. Motivated by these findings, we propose RobMAS, a two stage spatiotemporal reinforcement learning framework for adaptive communication recovery under agent dropout. In the spatial stage, graph theoretic rewards guide topology reconstruction toward sparse and connected communication DAGs. In the temporal stage, the Temporal State Recurrent Adapter (TSRA) maintains task relevant recovery history through recurrent node states, allowing previous recovery information to condition subsequent topology decisions. Together, spatial structural guidance and temporal recovery states are integrated within a unified recovery MDP, forming a closed loop self healing process for disrupted communication. Extensive experiments across three distinct domains demonstrate that, under a 40% random agent dropout rate, RobMAS maintains task accuracy close to its no dropout performance, with only 0.8%–3.0% degradation under static and dynamic dropout settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.