acceptodds
Under review as a conference paper at ICLR 2027

OfficeBench-Pro: Benchmarking Agents for Professional Office Work

Abstract

Professional Office work requires factual accuracy, analytical soundness, and aesthetic quality to jointly serve a workplace objective, yet existing benchmarks only partially capture their interdependence. We introduce , comprising expert-authored and synthesized tasks across professional domains. Agents create or revise Office deliverables from heterogeneous sources, with native structures essential to analysis or delivery. Our synthesis framework combines source-grounded design, execution-guided evolution, and independent task and rubric review. We adapt Agent-as-a-Judge to trace source evidence, test required native behavior, and visually inspect rendered outputs. Across six models under a shared execution protocol, the highest overall scores are 51.3% on expert tasks and 50.8% on synthesized tasks; no model leads all three quality dimensions in either collection. Qualitative analyses of archived submissions reveal unsupported premises propagated across deliverables, broken calculation dependencies, and visual organization that obscures analytical relationships.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.