World Model-Induced Dual-View Inconsistency for GUI Agent Trajectory Evaluation
Abstract
Benefiting from large-scale, high-quality interaction trajectories, GUI agents have achieved substantial progress in perception, planning, and interaction. As trajectory synthesis and automated collection scale up, data quantity is becoming less of a bottleneck, while trajectory quality is increasingly important. However, large-scale trajectory collections often contain substantial redundancy and quality variation, making reliable supervision difficult to identify.Trajectory evaluation is further complicated by rich interface and state information and unexpected factors during execution. Existing filtering methods mainly assess realized state transitions and cannot effectively distinguish action quality from deviations in execution outcomes, potentially discarding valuable supervision.Since identifying whether a realized transition matches the expected outcome requires predicting future states, we introduce WM-IncPRM, a world-model-guided inconsistency-aware process reward framework for GUI trajectory evaluation. WM-IncPRM combines a world model with PRM to predict future semantic transitions and compare them with realized transitions. Reliable and supported prediction–reality discrepancies are used to calibrate step-level quality, which is further aggregated with recoverable supervision coverage to classify trajectories into Keep, Repair, or Reject. Experiments demonstrate that WM-IncPRM improves trajectory filtering quality and downstream performance on AndroidWorld.We will release the source code on GitHub in the future.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.