FATE: First Agent Teams Exam
Abstract
Real-world professional work requires collaboration among coworkers with distinct responsibilities, information, and permissions, yet existing agent benchmarks provide limited coverage of this setting. We introduce First Agent Teams Exam (FATE), a benchmark for evaluating agent teams on realistic professional tasks. Constructed by 32 domain professionals, FATE comprises 112 tasks across 14 domains and 129 occupations, covering both unimodal and multimodal tasks. Agents use tools and communicate under role-specific resource and permission boundaries to produce intermediate artifacts and final deliverables. We evaluate eight LLMs and compare four communication protocols. Current agent teams achieve a best final evaluation score of only 42.0, with limited collaboration gains at substantial computational cost. These gains vary with communication protocols and execution parallelism, while coordination and artifact delivery remain major bottlenecks to reliable teamwork. We hope FATE will provide a foundation for systematically evaluating agent teams and advancing reliable agent collaboration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.