ConScore-VG: Consistency-Score-Guided Test-Time Reinforcement Learning for Visual Grounding
Abstract
Visual grounding requires models to localize image regions according to language instructions and is a fundamental capability of vision-language models. Recent approaches based on supervised fine-tuning or reinforcement learning have achieved strong performance, but adapting them to new scenarios without additional annotations remains challenging. Test-time reinforcement learning (TTRL) offers a label-free alternative, yet extending it to visual grounding is challenging because continuous coordinate outputs are not naturally compatible with standard consensus-based pseudo-labeling, and spatial consensus among predictions does not necessarily imply correct localization. To address these challenges, we propose ConScore-VG, a consistency-guided TTRL framework that combines spatial and semantic consistency scores for reliable pseudo-label selection. The Spatial Consistency Score measures how well each candidate localization agrees with the overall spatial consensus, while the Semantic Consistency Score evaluates region-instruction alignment through visual information gain. Extensive experiments on referring expression comprehension (REC) and GUI grounding benchmarks demonstrate that ConScore-VG consistently improves grounding performance across both settings. Notably, ConScore-VG yields consistent gains on REC with only 256 unlabeled samples and improves over the baseline by 6.68% on ScreenSpot-v2. These results demonstrate the potential of TTRL for visual grounding and establish ConScore-VG as a data-efficient approach for adapting vision-language models to downstream grounding tasks without labels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.