acceptodds
Under review as a conference paper at ICLR 2027

CoMateEval: Benchmarking LLMs Across Team Roles in Dynamic Multi-Party Collaboration

Abstract

Large language models are increasingly integrated into everyday workflows as interactive assistants. However, real-world work is often carried out in teams rather than through isolated human–AI interactions. Recent studies have begun to evaluate LLMs in collaborative settings, but existing formulations of collaboration are often constrained to specific tasks, roles, or interaction paradigms, leaving broader team responsibilities and dynamic multi-party collaboration insufficiently characterized. To address these limitations, we introduce CoMateEval, a benchmark grounded in Belbin team role theory for evaluating LLMs across diverse team responsibilities in dynamic multi-party collaboration. CoMateEval covers 9 team roles through 18 collaborative scenarios and instantiates them into 370 evaluation cases spanning diverse forms of team contribution. We further design CoMateSim, a multi-party collaboration simulator supporting constrained adaptive interaction, where predefined event structures and branch conditions govern scenario progression while interaction trajectories evolve with model behavior. Model performance is assessed by how well it fulfills its assigned responsibilities. Experiments on frontier LLMs show that the strongest model achieves an average rubric pass rate of 45.47% across the selected scenarios, while different models exhibit distinct strengths across team responsibilities. These results identify substantial gaps in fulfilling the tested responsibilities and motivate further work toward reliable AI teammates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.