WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing
Abstract
Enterprise agents often need several kinds of knowledge at once: documents for narrative facts, tables for calculations, and dependency graphs for file relationships. Existing benchmarks usually test retrieval or tool use, but do not separate choosing the right knowledge form from using it correctly. We introduce WorkSurface-Bench, which evaluates this choice as surface routing. The benchmark projects five persona-scoped Workspace-Bench-Lite workspaces onto document, table, and graph surfaces and contains 1,151 atomic tasks. Its reference answers are auditable: table answers come from executed DuckDB queries, document answers are checked against verbatim spans, and graph answers derive from source annotations. We evaluate four backbones under six agent settings, yielding 27,624 retained trajectories with no protocol errors. Gold-constrained agents reach 98.7–99.8 Route F1, yet Answer remains 56.1–75.3%. A matched intervention shows that surface hints improve Answer for three of four models, whereas removing irrelevant tools mainly improves routing and efficiency. In an independent three-annotator audit, all 200 sampled tasks pass all six criteria by majority vote. We release the construction pipeline, scoring code, and agent harness at https://anonymous.4open.science/r/WorkSurface-Bench-Anonymous-D097.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.