AssetArena: A Benchmark for Agentic Workflows in Industrial Asset Operations
Abstract
Industrial asset operations require agents to combine condition assessment, diagnosis, evidence retrieval, and operational action within a single stateful workflow. However, evaluations remain fragmented, as deployed agents typically cover only part of this workflow, while industrial benchmarks largely assess isolated capabilities or static knowledge rather than end-to-end execution. We introduce AssetArena}, an executable benchmark for agentic workflows in industrial asset operations. AssetArena provides a unified MCP-native environment that connects heterogeneous industrial evidence, analytical tools, and operational actions, with scenarios designed so that relevant information must be acquired through interaction. Its evidence-first construction process establishes reference outcomes independently of the agent and records replayable executions for reproducible evaluation. The benchmark contains 231 human-authored, machine-verified scenarios and 85 tools across six MCP servers. We evaluate eight frontier language-model configurations under a common agent framework and further study the effects of domain skills, agent decomposition, tool-access topology, and execution harness. Models achieve 16.5–29.9% Pass@1 across the benchmark, indicating substantial headroom even for frontier models. Controlled comparisons show that agent design and execution infrastructure affect task success and interaction cost: across 280 paired runs, Stirrup agent achieves 31.1% Pass@1 versus 25% for OpenCode. AssetArena provides a common and extensible environment for evaluating model capability and agent design on challenging industrial asset-operation workflows. Reproducibility materials are available at: https://anonymous.4open.science/r/AssetArena-ICLR/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.