acceptodds
Under review as a conference paper at ICLR 2027

JustJudge: Better Evidence Beats More Machinery in Judging GUI Agents

Abstract

Deciding from a recorded trajectory whether a GUI agent completed its task underpins how these agents are evaluated, how their training data is curated and how they are rewarded. To make that decision reliable, recent work adds machinery, such as critic pipelines and trained reward models. We show that the binding constraint is what the judge is shown. A trajectory is testimony as much as evidence: the agent under judgment writes part of the record, and a judge handed that record inherits the agent's verdict on itself. Prior judges warn the model about the record or discard the agent's text, and added machinery re-reads the same record. JustJudge instead sanitizes the record: it reduces each step to the action issued and a claim-free statement of what it was for, and masks recognized self-reported outcomes. The judge checks this step log against a fixed budget of screenshots anchored at the final state. The judge is a text-only extraction pass and one verification call to a vision-language model, with no training and no environment access. On OSReward, JustJudge with an open 27B model reaches 89.8% accuracy, level with the best proprietary judges the benchmark reports, and leads every judge on the hard split. The gain comes from the evidence, not the model: shown the sanitized record, a 9B model overtakes the reward model trained from it on that split. On the same backbone, published multi-call procedures cost up to 16 times as much and score 8 to 21 points lower. Downstream, the same judge produces stronger agents on OSWorld-Verified, as a filter on fine-tuning data and as the reward for online reinforcement learning. Reliable judging is a question of evidence, not of machinery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.