acceptodds
Under review as a conference paper at ICLR 2027

Learn From Move: Can GUI Agents Learn How Their Environments Work?

Abstract

GUI benchmarks increasingly emphasize longer workflows, but task horizon alone does not capture whether agents can learn how an environment responds to their actions. We introduce LatentGUIWorld, a benchmark of six tasks in three matched pairs organized by the evidence needed for control: uses prior knowledge and available observations, requires interaction that directly reveals hidden semantics, and identifies dynamics from action–response relationships. Frontier models struggle especially on tasks. We introduce MOVE, which separates movement from button events and exposes intermediate visual feedback, together with demonstrations of iterative exploration and control. Ablations show that both action granularity and multi-step demonstrations improve performance. After SFT, LatentLearner-9B achieves 50.0% overall success on LatentGUIWorld, comparable to several frontier models. Joint RL further raises egocentric Drag and Ten-Choice success to 92.0% and 75.3%. Behavioral analysis reveals multi-step exploration. In egocentric Drag, RL increases Pearson correlation between error-normalized commands and ideal gain compensation from 0.610 to 0.803. Together, these behaviors show that models learn to explore through interaction and adapt their control to environment dynamics. The same format improves GUI grounding after RL, reaching 67.37% mean accuracy across five benchmarks versus 66.72% for conventional RL.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.