OfficeBench-Pro: Benchmarking Agents for Professional Office Work
Abstract
Professional Office work requires factual accuracy, analytical soundness, and aesthetic quality to jointly serve a workplace objective, yet existing benchmarks only partially capture their interdependence. We introduce , comprising expert-authored and synthesized tasks across professional domains. Agents create or revise Office deliverables from heterogeneous sources, with native structures essential to analysis or delivery. Our synthesis framework combines source-grounded design, execution-guided evolution, and independent task and rubric review. We adapt Agent-as-a-Judge to trace source evidence, test required native behavior, and visually inspect rendered outputs. Across six models under a shared execution protocol, the highest overall scores are 51.3% on expert tasks and 50.8% on synthesized tasks; no model leads all three quality dimensions in either collection. Qualitative analyses of archived submissions reveal unsupported premises propagated across deliverables, broken calculation dependencies, and visual organization that obscures analytical relationships.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.