Group-Relative Self-Verified Policy Optimization for Label-Free GUI Grounding
Abstract
Reinforcement learning for GUI grounding commonly relies on binary rewards that indicate whether a predicted click falls inside an annotated target region. Such rewards become uninformative when all rollouts in a group share the same outcome, and they also collapse near misses and distant errors into the same supervision. Our analysis further reveals that many near misses are produced with high confidence, suggesting that models often identify the approximate target region but lack sufficient clicking precision. These observations suggest that the missing supervision is not necessarily additional labels, but a graded signal that can exploit the model's own approximate localization. Motivated by this, we propose Group-Relative Self-Verified Policy Optimization (GSVPO), which converts local re-observation and cross-trajectory self-verification into a continuous, label-free grounding reward. GSVPO samples multiple global click proposals, re-examines each proposal under an entropy-adaptive magnified view, and maps the refined predictions back to the original coordinate space. Candidates are then evaluated according to their agreement with the ensemble of refined predictions across trajectories, preserving graded spatial distinctions without requiring ground-truth coordinates. The resulting signal provides a unified interface for training-free candidate reweighting, label-free test-time reinforcement learning, and reward augmentation for standard GRPO. Across six GUI grounding benchmarks and six vision-language models, training-free GSVPO improves 32 of 36 model–benchmark pairs, with an average gain of 6.03 percentage points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.