Weakly-Supervised Cooperation in Decentralized LLM Multi-Agent Systems
Abstract
Existing state-of-the-art methods for learning inter-agent communication topologies in LLM multi-agent systems mostly learn through task outcomes. However, outcome-based rewards have significant challenges, such as requirement of correct task verification, a calibration phase, and also failing to adapt to changes in environment or team dynamics. In this work, we show that explicit task outcome signals are not necessary to learn effective communication topologies for cooperation. We propose a novel weakly-supervised method, MERIT-MAS, predicting dynamic and directed agent-to-agent communication topologies, only from agent-generated content. MERIT-MAS quantifies goal-conditioned relevance and novelty between agent pairs via cross-encoder rank scoring, and trains a topology network without using any task outcomes. Across established benchmarks and LLMs, our method achieves comparable performance (yielding overlapping 95% Wilson confidence intervals) with outcome-based SOTA baselines while spending 15% lesser agent-LLM tokens. Since MERIT-MAS does not need a calibration phase, it shows robustness under structural changes and scales seamlessly to larger agent counts, where methods that cache a single topology per episode collapse.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.