acceptodds
Under review as a conference paper at ICLR 2027

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic Tasks in Real-World Professional Fields

Abstract

Recent advances in AI agents have enabled increasingly complex interactions with real-world environments, yet existing GUI benchmarks largely focus on general-purpose software and short-horizon tasks. It therefore remains unclear whether modern agents can autonomously operate domain-specific professional software to complete long-horizon, economically valuable workflows end to end. To address this gap, we introduce Workflow-GYM, a benchmark for evaluating GUI agents on realistic professional workflows across diverse domains and specialized software environments. Extensive experiments show that even the strongest models achieve only slightly above 30% success rates, revealing substantial limitations in current GUI agents. Further analysis identifies several recurring failure modes, including workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. These results highlight long-horizon workflow consistency and reliable interaction with specialized software as key challenges for the next generation of GUI agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.