Counterfactual Evidence-Aware Credit with Matched Controls for Single-Step Reinforcement Learning of GUI Visual Grounding
Abstract
A GUI grounding policy can click the right location for the wrong visual reason. A binary click reward cannot tell whether the policy relies on the referred interface element or on a shortcut, and editing only the target cannot tell either, because a click can move after any image edit. We introduce counterfactual evidence-aware credit with matched controls (CEAC), which applies the same edit to the target and to same-size control regions outside it and turns the difference in the policy's response into training credit. We measure the matched gap: how much more the click moves when the target is edited than when the controls are edited. Compared with five other training objectives, CEAC increases the matched gap on held-out mobile screens by 15% to 22%, with no statistically significant accuracy loss. On PC/Web, CEAC also has the largest gap, and its lead over three of the five other objectives is significant after multiple-comparison correction. Mask and blur edits give the same gain, which comes from a larger response to target edits while control responses stay flat. Matched controls can thus train a policy to rely on evidence at the target location, which the task reward alone does not require.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.