TutorialGUI: Towards Generalizable GUI Agents via Tutorial-Driven Skill Learning
Abstract
GUI agents have made rapid progress in operating general-purpose software, yet they still face out-of-distribution (OOD) generalization challenges when deployed on unfamiliar specialized applications, where insufficient application-specific knowledge and operational experience limit reliable task completion. Scientific software exemplifies this challenge. However, adapting agents through per-application data collection and fine-tuning is costly, limiting the scalability of this approach to long-tail applications. Official tutorials offer readily available domain knowledge and procedural examples, but they are primarily written for humans, making them difficult for agents to use directly, and their coverage is limited. To address these limitations, we introduce TutorialGUI, a training-free framework that transforms official tutorials into executable and extensible Atomic Skills to improve GUI agents’ generalization to long-tail applications. At its core is a Replay–Explore loop: Replay converts tutorial procedures into validated skills through actual GUI execution and outcome verification; Explore builds on these validated skills to discover and verify related capabilities, extending coverage beyond the original tutorials. The resulting skills can be retrieved and reused by different GUI agents without updating model parameters. On scientific software tasks in ScienceBoard, TutorialGUI improves task success rates by an average of 14.285 percentage points across multiple models compared with no-tutorial baselines, demonstrating the effectiveness of transforming human-oriented tutorials into verifiable and extensible capabilities for GUI interaction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.