EarlyKnowBench: When Does a Gui Agent Know What the User Wants, and Does It Keep Knowing?
Abstract
A proactive mobile GUI agent must infer user intent from partial interaction trajectories early enough to intervene, while reliably sustaining that prediction over time. However, existing GUI benchmarks evaluate only a single terminal state of a completed trajectory, failing to distinguish a model that discovers and retains the correct goal early from one that reaches it too late to act. To address this limitation, we introduce EarlyKnowBench, a benchmark that probes the latent session goal at every observation step and scores the prediction sequence against an annotated full evidence support time (t*), defined as the earliest step where observed screenshots and executed actions fully specify the goal. Grounding t* in trajectory evidence rather than model-dependent outputs enables standardized, fair comparisons of early intent recognition and temporal retention across sessions and models. We annotate t* across 1,503 sessions comprising 29,630 interaction steps from FingerTip-20K and AndroidControl, evaluating 12 multimodal models across every step. Empirical results reveal that current models struggle to infer user goals prior to t*, and even when successful early on, they systematically fail to retain these predictions across subsequent steps. Furthermore, downstream experiments demonstrate that accurate goal predictions substantially boost action execution success rates. These findings indicate that future proactive agents must look beyond raw scaling to focus explicitly on early intent recognition, temporal prediction retention, and downstream execution alignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.