VISTA: Verifiable Self-Curriculum for Training GUI Agents
Abstract
Training GUI agents requires scalable supervision for long-horizon interactions in stateful environments, where success is often difficult to specify and verify. Existing approaches rely on human annotations, large teacher models, or manually specified evaluators. We introduce VISTA, a lightweight framework for verifiable self-curriculum learning that uses a small VLM to generate training tasks and their executable evaluators. Conditioned on the agent’s current capabilities, the model synthesizes tasks, environment configurations, and evaluators, retaining task–evaluator pairs that pass execution-based verification. The verified tasks support agentic reinforcement learning, and the updated agent informs subsequent task synthesis, forming a self-improvement loop. Our models VISTA-8B and VISTA-9B post-trained on top of Qwen3-VL-8B and Qwen3.5-9B show up to 4.7% and 4.3% absolute improvement of resolve rate on OSWorld-verified and WAA-V2 over their respective base models. Training with self-curriculum also improves performance on OSWorld-G and ScreenSpot-Pro even without grounding-specific supervision and reduces the average number of interaction steps. These results show that GUI agents can be improved through self-generated, executable supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.