MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments
Abstract
Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored. To address this gap, we introduce **MECoBench**, a multimodal embodied cooperation benchmark with a controlled evaluation platform spanning two distinct task domains and diverse task and team settings. Through extensive experiments across various MLLMs, we summarize three key findings: (i) Most models benefit from collaboration, but individual capability alone does not determine collaborative performance. (ii) Communication is essential for collaboration gains, with reliable information sharing and coordinated execution distinguishing effective collaboration from failure. (iii) Appropriate team size, organization, and information sharing can further improve collaborative performance. Generally, MECoBench provides a systematic testbed for understanding the mechanisms and limits of multimodal embodied collaboration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.