TRACK: Autonomous Exploration and Task-Progress-Conditioned Knowledge Utilization for Mobile GUI Agent
Abstract
Offline knowledge acquisition and online knowledge utilization pipeline provides application-specific interaction knowledge for long-horizon mobile GUI tasks. However, existing methods do not sufficiently align how knowledge is acquired and reused with the evolving information needs of long-horizon GUI decision-making, often resulting in incomplete guidance or knowledge that is mismatched to the current task progress. Hence, we propose TRACK, a GUI Agent framework that constructs a multi-granularity knowledge base through Offline Exploration and Knowledge Synthesis, and adaptively injects the knowledge required at each step through Task-Progress-Conditioned Knowledge Utilization. During the offline stage, TRACK performs traceable and verifiable autonomous branch exploration around the major functionalities of each application. It then synthesizes the resulting exploration branches into complementary task-path knowledge for what to do, element knowledge for where to act, and skill knowledge for how to execute, forming a multi-granularity knowledge base. During online execution, we use the knowledge base to locate the current task progress from historical executions and to route the knowledge applicable to each decision step. By coupling multi-granularity knowledge base with offline autonomous exploration and progress-conditioned online utilization, TRACK provides complementary and progress-aligned guidance throughout long-horizon interactions. Experiments on AndroidWorld and DroidTask show that TRACK achieves strong task completion performance with marginally higher execution efficiency on Qwen3.5-Plus, while its constructed knowledge base further improves the success rates of integrated agents by up to 11.63% and simultaneously reduces their execution steps.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.