Beyond Coordinates: Semantic Reward for GUI Agent Post-Training
Abstract
Graphical User Interface (GUI) agents have made substantial progress in automating tasks with vision-language models. Reinforcement learning has further improved GUI grounding by optimizing policies with coordinate-based rewards, which provide direct supervision for spatial localization. However, GUI interaction data also contains rich semantic information about the intended operation, target element, and expected outcome, which cannot be fully captured by geometric objectives alone. To exploit this complementary information, we introduce semantic rewards for GUI agent post-training. During reinforcement learning, candidate actions are represented as natural-language descriptions of the intended GUI operations, and a VLM evaluates their semantic consistency with the target interaction. We further propose Self-Reward RFT, where the semantic evaluator is initialized from the policy and periodically updated with the latest policy weights while remaining frozen between updates. This design avoids maintaining a separately trained reward model while providing an adaptive semantic training signal. Extensive experiments across multiple GUI grounding benchmarks demonstrate the effectiveness of semantic supervision across different policy backbones. Using only 6k training samples, our framework with Self-Reward RFT improves MAI-UI-2B from 56.7% to 59.0% on the challenging ScreenSpot-Pro benchmark, achieving strong performance among lightweight GUI models while remaining competitive with several larger models. Further analysis shows that semantic and coordinate-based supervision capture complementary aspects of GUI action correctness, highlighting the value of incorporating semantic feedback into GUI agent post-training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.