acceptodds
Under review as a conference paper at ICLR 2027

Where Does A Predictive GUI Verifier Lose To A Discriminative One? A Controlled Diagnosis

Abstract

Step-level verification asks whether an observed GUI screen is compatible with the action claimed to have produced it, the goal-free check that retry and recovery rely on. A predictive verifier forecasts the successor with a latent world model and compares the forecast with the screen; a discriminative verifier reads the action and both screens jointly. We compare the two at matched data, negatives, backbone and training budget on a pre-registered benchmark of recorded transitions. The predictive verifier is worse by 6.53 AUC points, by 6.30 on a test set built and scored once after every model was frozen, and by a wider margin at low false-positive rates. The gap is not a labelling artefact: two independent annotators find a minority of constructed negatives still compatible, and on the negatives both confirm the gap is essentially unchanged. A ladder of controlled variants between the two verifiers locates it. The four losses that align the forecast with the observed screen move aggregate AUC by only 0.19 points. Comparing the screens after encoding them apart costs 4.93; in-distribution, neither keeping every patch nor matching the comparison head recovers more than a point of it, which leaves fusion timing. Gains that narrow the gap do not survive a change of distribution, and both verifiers refuse genuine interface anomalies more often than ordinary transitions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.