Learning When to Observe: Visual Readiness for GUI Agents
Abstract
GUI agents must decide when to observe an interface again after taking an action. Existing systems often use fixed delays or global visual stability, conflating observation timing with action execution timing. We show that these are distinct control factors. Using identical-seed counterfactual replay with a frozen vision-language actor, we independently vary observation maturity and execution delay. Across two MiniWoB task families, early observations produce incorrect coordinates that cannot be repaired by delaying execution, while mature observations still require sufficient execution readiness. Matched no-delay controls isolate these timing effects from grounding failures. We further construct static-pending and moving-ready interfaces, where global pixel stability releases too early or waits unnecessarily. Motivated by these findings, we formulate post-action control as instruction-conditioned visual readiness and train a binary controller from counterfactual frozen-actor outcomes. On 80 unseen episodes, the controller matches a validation-locked 0.5-second fixed wait at 83.75% success, while reducing mean observation time from 0.500 to 0.132 seconds and mean actual execution time from 1.448 to 1.066 seconds. These results establish post-action observation timing as a distinct control problem and show that learning when to observe can preserve task success while reducing latency under a fixed actor-query budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.