OpenMAS-GCom: A Diagnostic Benchmark for Graph-Enhanced Multi-Agent Systems
Abstract
Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through structured communication and role assignment, which determine information exchange and collaboration patterns. As G-MAS develops toward complex collaborative reasoning, comprehensive evaluation becomes increasingly important. However, existing benchmarks provide limited diagnostic evaluation of collaboration. They mainly evaluate final task performance, providing limited insights into how G-MAS approaches handle information integration, erroneous information propagation, and execution recovery during collaboration. To address this limitation, we propose OpenMAS-GCom, a comprehensive diagnostic benchmark for G-MAS evaluation. Specifically, (1) From the evaluation perspective, OpenMAS-GCom introduces a controlled intervention framework that represents G-MAS through collaboration units, communication links, shared intermediate information, and execution strategies. It enables systematic analysis of organizational components while keeping tasks, models, and resource settings fixed. (2) From the task perspective, OpenMAS-GCom introduces G-MAS-Complex, a new benchmark containing 400 multi-document tasks that require information integration, conflict resolution, and structured answer generation. OpenMAS-GCom covers 29 datasets across six domains and evaluates 17 representative configurations, including single-agent, ordinary multi-agent, and graph-enhanced multi-agent systems. Extensive experiments reveal that graph enhancement does not uniformly improve G-MAS performance. On G-MAS-Complex, the advantage of graph-enhanced approaches is mainly achieved by a small subset of configurations, with graph-based methods showing large performance gaps on increasingly difficult collaboration tasks. Moreover, methods with comparable initial accuracy exhibit substantially different robustness under corrupted information, revealing hidden differences in collaboration reliability beyond final task scores.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.