acceptodds
Under review as a conference paper at ICLR 2027

Traceverifier: Evidence-Linked Claim Graphs For Gui Agent Trajectory Verification

Abstract

Reliable verification is essential for evaluating GUI agents, curating training trajectories, and producing learning signals. Yet observable-trajectory verification remains difficult: plausible final states can conceal instruction-violating execution, while long trajectories can obscure decisive evidence. A single task label further hides the evidence behind a decision and whether downstream requirements were blocked by earlier errors. We formulate a new task, Structured Trajectory Verification, which maps a task instruction and an observed GUI trajectory to a structured verification report rather than a single outcome label. To address this task, we present TraceVerifier, which resolves instruction-derived claims using localized evidence, semantics-matched operators, and assessment dependencies, producing an evidence-grounded report that distinguishes direct failures from dependency-blocked requirements. We also construct TraceVerifierBench, comprising 384 mobile GUI trajectories from 85 applications, with outcome labels and two fine-grained diagnostic pilots. Across three task-outcome evaluations, TraceVerifier obtains the highest accuracy and pass precision among the evaluated verifier methods, exceeding the strongest method baseline on each benchmark by 2.85–12.71 accuracy points. Beyond final-label performance, its structured reports achieve 82.27% evidence-localization and 81.59% immediate-blocker , demonstrating measurable evidence grounding and failure attribution. Finally, TraceVerifier's structured outputs serve as fine-grained supervision: Process-Trace Distillation yields a best observed Qwen3.5-9B accuracy of 85.66%, within 0.82 points of the Qwen3.5-27B teacher.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.