acceptodds
Under review as a conference paper at ICLR 2027

VRC: Visual Reward Consistency in GUI-Agent Reinforcement Learning

Abstract

Visual reward models can classify recorded trajectories accurately yet provide weak feedback for policy learning. We examine this gap in graphical user interface (GUI) reinforcement learning using paired visual evidence on BrowserGym/MiniWoB. Four data-matched verifier objectives compare binary cross-entropy with invariance and failure-sensitivity losses. A separate policy experiment compares a BCE verifier's single-view reward with five-view soft-min scoring from the combined-loss verifier. Across three seeds, sensitivity supervision achieved the highest validation AUROC, while BCE had the lowest banner inconsistency. On 2,160 evaluation episodes across 24 task families, BCE reward achieved 33.06% success and the soft-min reward achieved 32.22%. Neither the original comparison nor a post-hoc crossed seed-task analysis resolved an improvement from soft-min. Most held-out successes came from one task family, and the soft-min reward showed weak separation between successful and failed policy trajectories. The policy experiment changes verifier training and aggregation together and excludes the sensitivity-only verifier. The observed result concerns this reward combination; the downstream value of sensitivity supervision remains open.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.