acceptodds
Under review as a conference paper at ICLR 2027

ProfessionalBench: Benchmarking AI Agents on Real-World Professional Workflows

Abstract

As AI agents handle files, use tools, and produce complex deliverables, evaluating their ability to perform professional work becomes increasingly important. Recent benchmarks increasingly adopt workflow-based tasks to improve task realism. However, workflows may undergo Workflow Compression during task construction: responsibilities that belong to the designated role are omitted or resolved in advance, preventing direct evaluation of whether agents can fulfill them and where they fall short. To capture real professional workflows, we introduce ProfessionalBench, spanning 30 subdomains across finance, law, and healthcare. Experts compare tasks against source workflows and use a Know–Act–Close (KAC) formulation to preserve interdependent responsibilities for professional judgment, execution, and closure. Expert-reviewed Dense Rubrics, aligned with KAC responsibilities, support hybrid verification of requirement completion and deliverability, while trajectory analysis examines the failure mechanisms behind unmet requirements. We evaluate seven frontier agents on ProfessionalBench. Results reveal that even the strongest agents achieve pass rates below 40%. Workflow Compression interventions show thatomitting or pre-resolving responsibilities increases mean scores, with larger gains from compressing Act (+4.1 points) and Close (+3.7) than Know (+2.2). Trajectory-based diagnosis finds failures across all three responsibilities, with the largest shares attributed to Act in finance, Know in law, and Close in healthcare. Together, these findings show that preserving role-specific responsibilities exposes substantial professional headroom and identifies where agents need to improve.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.