From Future to Present: Utility-Routed Future Distillation for Mobile Agent
Abstract
Mobile GUI agents rely on current screens and interaction history, yet anticipating the state a task should reach next offers additional guidance for action decisions. Although logged trajectories naturally provide future screens, these observations are unavailable at deployment and are not uniformly beneficial for action prediction. We propose Utility-Routed Future Distillation (URFD), which transfers useful future representations into a current-only GUI policy without additional annotations or future-image generation. Task- and Action-Conditioned Successor Initialization (TASI) combines dual-view action prediction with future-selection tasks to establish a future-aware initialization. URFD distills future-conditioned representations into a compact latent slot within a current-only policy. A frozen dual-view teacher quantifies future utility through the gain in demonstrated-action log-likelihood, jointly governing the interpolation of hidden-state targets and the weighting of policy distillation across current- and future-conditioned views. This utility-guided formulation mitigates interference from unhelpful future observations during distillation. The experimental results show that URFDimproves average action match by 2.26 percentage points across two offline benchmarks and online task success by 3.5 points on AndroidWorld. Comprehensive analyses examine the utility of future supervision, strategies for representation alignment, and their impact on action decisions. These results demonstrate that URFD enables models to approximate hidden representations of real future states and improve action prediction accuracy without observing real future states at deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.