acceptodds
Under review as a conference paper at ICLR 2027

VeriCUA: Learning Proactive Verification for Reliable GUI Agent Rewards

Abstract

Reinforcement learning for computer-use agents increasingly relies on proactive reward agents that interact with the post-execution GUI environment to verify task completion beyond what recorded screenshots reveal. The reliability of such verification hinges on two coupled decisions, namely what to check and when to stop. Existing methods leave these decisions to prompted foundation models, yet neither model scale nor few-shot examples help, as prompted Judges ranging from 4B models to GPT-5.4 all reach between 68% and 70% macro-average agreement. These decisions are hard to prompt because what a check yields and when the evidence suffices depend on the environment rather than the task, and become apparent only once verification is executed. Learning them from the environment is nevertheless difficult, since verification starts from the Actor's terminal state, which cannot be faithfully replicated during training, so each state admits a single attempt. We therefore introduce VeriCUA, which makes each single attempt informative and compares attempts across terminal states. A trainable Judge decides what to check and when to stop, while a frozen Probe executes the checks in the live environment. The Judge is trained with online reinforcement learning on one verification trajectory per terminal state, with a reward reflecting both verdict correctness and Probe usage and returns normalized across terminal states. Trained only on OSWorld and evaluated without further training on MobileWorld and WebArena-Verified, VeriCUA attains the highest reward agreement on all three benchmarks and lies on the agreement–cost Pareto frontier among proactive verifiers, outperforming few-shot GPT-5.4 by 5.36 macro-average points with 47.6%–59.8% fewer Probe tokens. As the reward for policy training, it raises task success by 3.21 points overall and 5.40 points on held-out multi-application tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.