acceptodds
Under review as a conference paper at ICLR 2027

VeriCUA: Verification as Agentic Exploration for Computer-Using Agents

Abstract

Computer-using agents (CUAs) follow natural-language instructions on desktops, mobile devices, and web browsers by reading and acting within the OS. Their evaluation, data curation, and reinforcement learning all require verifying at volume whether each trajectory fulfilled its instruction. Hand-written checkers stop at benchmark scale and human annotation cannot keep pace, so the field increasingly turns to VLM judges and automatically generated verifiers, each with its own reliability problem, from the false positives of judges to checks never grounded in the environment. Both rest on a static verifier that sees only partial information, while the evidence of success may be scattered far beyond that view. To overcome this bottleneck, we present VeriCUA, a new paradigm for verifying CUA tasks in which the verifier adapts to and actively explores its environment for evidence of task completion. We equip it with an adaptive, dynamic action space combining code and GUI to navigate recorded trajectories and online environments alike. VeriCUA decomposes the task into a rubric of end-state requirements, plans and explores at test time to gather evidence for each, and types each piece by how it was obtained. Only attested effects, never the agent’s own assertions, satisfy a requirement, and a deterministic arbiter then derives the verdict from this evidence. Experiments show VeriCUA to be a promising solution across online and offline settings. On recorded trajectories across multiple testbeds and benchmarks, it surpasses typical judging methods and improves markedly on hard cases. On online environments, it certifies far fewer failed runs as successes than the agent’s own completion claim and agrees more often with official evaluators. Analyses trace the gain to the action space, find that it grows with a stronger backbone, and show where recorded evidence runs out. Its evidence-backed verdicts are precise enough for scalable generation of CUA reward signals.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.