DeliverBench: Benchmarking AI Agents on Delivering Realistic Multi-Stage Professional Projects
Abstract
As LLM agents grow more capable, delegation is expanding from individual assistance to end-to-end professional project delivery, with potential to support enterprise automation and even one-person companies. Some existing benchmarks cover professional tasks and extended workflows, but inadequately assess how projects evolve and whether their final multi-file deliverables meet acceptance criteria. Thus, we introduce DeliverBench, which takes a multi-stage professional project as its evaluation unit. Reconstructed from real-world projects with domain experts, it includes 11 projects across seven domains and 150 interdependent subtasks. Agents work in a persistent workspace, incorporating new materials and client feedback released along a project timeline while reusing, revising, and integrating their own earlier artifacts. Expert-written rubrics assess checkpoint submissions and evaluate final packages for deliverable quality, cross-file consistency, and package acceptance. Across seven model–harness systems, we find that mean final-delivery scores range from 48.1 to 70.0 out of 100, and subtask pass rates at a 90-point threshold range from 13.3% to 50.0%. Every system's mean final-delivery score falls below its mean process score, revealing a gap between local task performance and project delivery. Correctness and usability are the weakest dimensions, while case studies highlight sustained quality and artifact repair. Runtime and token use vary widely, with greater expenditure not consistently associated with better delivery. These findings indicate that task-level scores alone cannot establish project readiness, motivating evaluation of agents’ ability to sustain professional work and deliver coherent, usable packages.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.