FORGE: Feedback-driven Optimization and Refinement for GUI Evolution
Abstract
GUI agents based on vision-language models have shown strong potential for general-purpose desktop automation, but their post-training remains constrained by the lack of scalable and reliably verifiable training data. Existing synthesis pipelines are largely static, making it difficult to adapt to new data quality issues and the evolving capabilities of trained models. We present FORGE, an end-to-end data flywheel that produces both high-quality supervised fine-tuning trajectories and verifiable reinforcement learning tasks. FORGE constructs diverse GUI tasks through scenario-behavior matching and pairs each task with an executable evaluator. Crucially, the pipeline evolves through its own feedback: recurring generation errors are distilled into reusable validation skills, while model failures on exploratory tasks inform subsequent data production. This closed loop progressively improves both data quality and task coverage. Using FORGE, we construct over 10,000 verifiable GUI tasks and evaluate the resulting models on OSWorld across two backbones. FORGE-8B and FORGE-4B achieve 51.2% and 40.4%, improving over their base models by 17.3% and 9.0%, respectively. We will release the dataset, code, and trained models upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.