GUIVista: Training GUI Agents from Tutorial Videos with Offline Capability Diagnosis
Abstract
Public software tutorial videos offer scalable supervision for GUI-agent mid-training, but their visual transitions often underdetermine the action that caused them. Even when these interactions are used for training, a failed action does not reveal whether the model chose incorrectly or struggled to express and execute the right action reliably. We present GUIVista, a framework that combines reliable trajectory construction with offline capability diagnosis. GUIVista revisits uncertain action proposals using transition evidence and action-family verification before filtering them into grounded, multi-turn training trajectories. It also evaluates checkpoints by asking them to select a successful action from plausible alternatives under the same GUI context, separating action understanding from open-ended execution. On an action-recovery evaluation, GUIVista improves overall action-type and full-action accuracy over prior video-based systems from 24.75% / 8.80% to 50.42% / 31.29%. Under the shared downstream SFT protocol, GUIVista mid-training improves end-to-end OSWorld AVG@4 from 18.40% for SFT-only to 21.75%. The offline diagnostic tracks post-SFT OSWorld performance across matched checkpoints, with Pearson and Spearman . These results show that GUIVista can turn noisy tutorial videos into useful training supervision and provide an efficient signal for comparing the capabilities acquired during mid-training. We will release all code for data construction and evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.