GCFSim: Learning Goal-Conditioned Functional Similarity for Mobile GUI Navigation
Abstract
GUI agents trained on large corpora of mobile-application trajectories can solve a wide range of user-requested tasks. A key challenge, however, is adapting such agents to a particular application without increasing inference latency. A promising approach is to explore the application offline, store application-specific knowledge, and retrieve relevant experience during navigation. This raises a fundamental question: which previously observed screen should the agent retrieve when deciding what to do next? We show that neither visual similarity nor goal progress alone is sufficient for this purpose. Two screens may look very different yet support equivalent useful actions for a particular task, while two states at the same distance from the goal may require different next actions. We formulate application-memory retrieval as goal-conditioned graph localization. Given the current screenshot and user goal, the objective is to retrieve a state from the explored application graph that provides useful guidance for the agent's next decision. For each goal, graph states are grouped according to the optimal actions they support and the goal-directed continuations reached through those actions. We use these functional relationships as supervision to learn GCFSim, a lightweight goal-conditioned visual-language representation . Although graph structure is used to construct training supervision, inference requires only the current screenshot and user goal. We evaluate GCFSim on a new cross-platform navigation corpus spanning 49 Android and HarmonyOS applications, with 7,903 observed UI states, 39,154 transitions, and 36,351 automatically generated navigation intents. The code and dataset are available anonymously for review and will be publicly released upon acceptance. Using the same application memories, functional supervision retrieves states that provide more useful navigation guidance than visual similarity, shortest-path distance supervision, and goal-conditioned value or Q representations. These results show that a useful retrieved state need not resemble the current screen or represent the same amount of progress toward the goal. What matters is whether the two states play the same functional role in completing the user's task.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.