acceptodds
Under review as a conference paper at ICLR 2027

Why Do Computer-Using Agents Fail? A Trace-Grounded Taxonomy and Its Reliability

Abstract

Computer-using agents (CUAs) are increasingly evaluated on graphical interfaces and live production websites, yet benchmarks typically report only a single terminal reward. This conflates failure origins across the model, agent scaffold, benchmark evaluator, execution harness, and task specification. We introduce a frozen CUA failure taxonomy—6 families and 18 categories—built by grounded-theory analysis over a -trace construction corpus, and apply it through a three-stage judging pipeline that decouples execution outcome from task success and attributes failures across system components. We evaluate on a disjoint held-out corpus of traces produced by a Harbor-based framework running four unmodified agent CLIs across two generations of OSWorld and ClawBench respectively. Category-level classification reaches Krippendorff's on the construction set and holds between and across the 26 held-out populations with and two blind annotators agree with the pipeline at –, against with each other over a validation set. The diagnostic infrastructure proves to be part of the measurement: trace normalization shifts attribution by percentage points, withholding screenshots raises the agent share by , and one revision of the judge's prompt moves it by roughly twenty—each exceeding the judge's own sampling noise.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.