ORGWorld: Benchmarking AI Agents as Collaborative Coworkers in Realistic Company Environments
Abstract
Can an agent carry out a full day of organizational work correctly without supervision? Such work requires an agent to infer obligations, reprioritize ongoing work as conditions change, and coordinate with colleagues under time constraints. Existing agent benchmarks provide limited evidence of whether agents can reliably fulfill organizational responsibilities under these combined demands. We introduce , a benchmark of 56 simulated company environments, each a full workday built around nine business systems, a company handbook, and 15 simulated colleagues. Each workday combines routine operational responsibilities with a cross-department project whose subtasks only colleagues can complete, requiring agents to carry company procedures through dependent steps and maintain progress across responsibilities as new events arrive. We combine expert review and solvability checks with programmatic verification of business outcomes, prohibited actions, and temporal constraints. A workday is solved only when all required responsibilities are completed and no prohibited action occurs. Across 12 frontier models, the strongest solves just 48.2% of workdays, while no other model exceeds 20%. These results reveal a substantial gap between current frontier agents and the reliability required to carry out organizational work autonomously.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.