acceptodds
Under review as a conference paper at ICLR 2027

DWBench: Evaluating AI Agents as Digital Workers over Enterprise Graph Data

Abstract

Enterprise knowledge work is broader than writing code: a digital worker must locate the right thread in a noisy inbox, reconcile facts scattered across emails, files, meetings, and chats, then produce a report, spreadsheet, or slide deck. We present DWBench, an agent-interface-agnostic benchmark for grounded retrieval and Office-artifact generation over ego-centric, permission-scoped enterprise graphs. We release two synthetic enterprise tenants (biotechnology and financial advisory) comprising 80 acting identities and 6,110 content units across email, chat, meetings, and 15–16 file types per tenant; 474 access-safe tasks spanning single-fact retrieval, multi-hop reasoning, Word, Excel, and PowerPoint production; content assertions linked to supporting records; and a common Model Context Protocol environment. We evaluate GitHub Copilot (GHCP), OpenCode, as well as a reference agent (RefAgent) across 13 model settings. The experiments comprise 15K executions across three harnesses, two tenants, and five task types using frontier models, including GPT-6-Astra, GPT-5.6-Sol, and Opus-5. They include paired-harness comparisons and model-scaling studies. Results show that mean assertion pass rate varies with the model, harness, and task type: Word generation attains mean assertion pass rates of 87.8–92.8%, whereas multi-hop search remains a bottleneck at 54.9–60.1%. GHCP achieves the highest mean assertion pass rate, exceeding OpenCode and RefAgent by 1.5 and 2.2 percentage points, respectively, while significant GHCP advantages over OpenCode in task completion rate at the 0.75 assertion threshold (a query is complete when at least 75% of its assertions are satisfied) are concentrated in specific models. Our findings provide reproducible diagnostics of tool-trace depth, model scaling, and evaluation robustness across enterprise environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.