acceptodds
Under review as a conference paper at ICLR 2027

Can Agents Compete as a Team? Benchmarking Long-Horizon Multi-Agent Collaboration with Realistic Contests

Abstract

Large language model (LLM)-based multi-agent systems (MASs) have demonstrated extraordinary performance in numerous tasks and domains, but their collaborative capability in challenging real-world tasks is still largely underexplored. To bridge this gap, we propose ContestWorld, a benchmark dedicated to evaluating agent collaboration with diverse team-based human contests. It consists of 321 contests from different years with 39 categories, such as MCM, ICPC, and ARML, spanning 9 domains, ranging from mathematics to computer science, business, and linguistic contests. ContestWorld is deployed on a carefully designed realistic contest simulator that closely aligns with human settings, with an action interface comprising 27 common actions, three optional memory actions, and contest-specific tools, filtered according to contest rules and protocol settings. For instance, in an ICPC contest, three teammates share a single compiler, so only one of them can use it at a time. In ContestWorld, agents must allocate tool calls within the limited turns, interacting with the teammates and environments to achieve the best performance. With ContestWorld, we have conducted a comprehensive evaluation of four primary baselines and six backbone LLMs. The best-performing GPT-6 model achieves only 63.4% on average, demonstrating the challenge of the proposed dataset. Our analysis shows that teams primarily benefit from improved task coverage and selection, while their remaining failures stem from weak verification and poor feedback integration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.