BIMArena: Benchmarking Domain-Specialized Computer-Use in Professional Building Modeling
Abstract
Computer-use agents are advancing rapidly; yet, existing benchmarks focus on general-purpose applications and do not fully reflect the complexity of real-world workflows. Moreover, they tend to prioritize breadth across applications over depth within professional workflows that span diverse functionalities and tightly coupled operations. In this work, we introduce BIMArena, a domain-specialized benchmark for evaluating the depth of computer-use agent capabilities in Building Information Modeling (BIM) authoring environments, a complex, knowledge-intensive professional domain that remains underexplored in current benchmarks. BIMArena comprises 409 tasks that progressively evaluate fundamental professional software operations, long-sequence workflows, and domain reasoning tasks. We further investigate whether external operational support through software documentation and expert-curated operational skills improves agent performance. Experiments are conducted in isolated real-computer environments using a programmatic evaluation framework. They reveal a pronounced capability gap: the strongest agent achieves a 98.4% success rate on fundamental software operations, yet only 26.6% on realistic, professional tasks. Our analysis shows that models strong on general computer-use benchmarks can still struggle with professional tasks beyond basic user interface operations. Operational support also becomes less effective with increasing task depth and can even be detrimental.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.