acceptodds
Under review as a conference paper at ICLR 2027

ORGWorld: Benchmarking AI Agents as Collaborative Coworkers in Realistic Company Environments

Abstract

Can an agent carry out a full day of organizational work correctly without supervision? Such work requires an agent to infer obligations, reprioritize ongoing work as conditions change, and coordinate with colleagues under time constraints. Existing agent benchmarks provide limited evidence of whether agents can reliably fulfill organizational responsibilities under these combined demands. We introduce , a benchmark of 56 simulated company environments, each a full workday built around nine business systems, a company handbook, and 15 simulated colleagues. Each workday combines routine operational responsibilities with a cross-department project whose subtasks only colleagues can complete, requiring agents to carry company procedures through dependent steps and maintain progress across responsibilities as new events arrive. We combine expert review and solvability checks with programmatic verification of business outcomes, prohibited actions, and temporal constraints. A workday is solved only when all required responsibilities are completed and no prohibited action occurs. Across 12 frontier models, the strongest solves just 48.2% of workdays, while no other model exceeds 20%. These results reveal a substantial gap between current frontier agents and the reliability required to carry out organizational work autonomously.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.