Agentic-WorkBench: Evaluating Multimodal Agents on Complex Visual Workflows
Abstract
Multimodal work often requires agents to acquire visual evidence, use it to guide successive actions, and verify the resulting artifacts or service states. Existing work-agent evaluations leave open how reliably visual information can support this full workflow. We introduce Agentic-WorkBench, a benchmark of 115 tasks across eleven work domains in controlled, stateful environments. The benchmark targets workflow-level visual intelligence, organized into visual grounding, sensemaking, and realization, with observations that may become available only during execution. Across nine model configurations, the strongest achieves 64.0% graded completion, solves 41.7% of tasks within three trials, and succeeds in all three on only 22.6%. The lowest success rates concentrate on visual sensemaking demands, and success varies substantially across work domains and objectives. Harness effects vary across models, and greater resource use does not consistently accompany higher success. These results identify a gap between partial progress and repeatable completion, motivating agents that maintain visual state and use feedback throughout practical work.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.