acceptodds
Under review as a conference paper at ICLR 2027

BIMArena: Benchmarking Domain-Specialized Computer-Use in Professional Building Modeling

Abstract

Computer-use agents are advancing rapidly; yet, existing benchmarks focus on general-purpose applications and do not fully reflect the complexity of real-world workflows. Moreover, they tend to prioritize breadth across applications over depth within professional workflows that span diverse functionalities and tightly coupled operations. In this work, we introduce BIMArena, a domain-specialized benchmark for evaluating the depth of computer-use agent capabilities in Building Information Modeling (BIM) authoring environments, a complex, knowledge-intensive professional domain that remains underexplored in current benchmarks. BIMArena comprises 409 tasks that progressively evaluate fundamental professional software operations, long-sequence workflows, and domain reasoning tasks. We further investigate whether external operational support through software documentation and expert-curated operational skills improves agent performance. Experiments are conducted in isolated real-computer environments using a programmatic evaluation framework. They reveal a pronounced capability gap: the strongest agent achieves a 98.4% success rate on fundamental software operations, yet only 26.6% on realistic, professional tasks. Our analysis shows that models strong on general computer-use benchmarks can still struggle with professional tasks beyond basic user interface operations. Operational support also becomes less effective with increasing task depth and can even be detrimental.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.