Exact Visual Grounding Strengthens Visual Information Utilization in Vision-Language Models
Abstract
Modern vision-language models (VLMs) achieve strong semantic reasoning yet remain unreliable on questions requiring precise visual evidence, such as exact counts, distances, or geometric relations. Broad multimodal pre-training provides rich semantic supervision, but rarely requires models to ground their predictions in exact visual evidence, leaving room for semantic and language priors to dominate when precise perception is required. To address this limitation, we introduce Exact Visual Grounding (EVG), a framework that provides explicit supervision for grounding answers in precisely controlled visual properties. EVG adopts an answer-first construction: it first samples a target answer, then renders an image whose relevant visual property exactly realizes that answer, and finally generates the corresponding question. This yields annotation-free supervision across 12 Color, Counting, and Geometric Measurement tasks, together with matched missing-object examples. Empirically, we find that training on these tasks consistently improves downstream multimodal performance. For example, on Qwen3.5-9B, EVG raises the macro-average across nine benchmarks from 60.06% to 63.90%, with similar gains across different model scales and backbones. To understand how EVG alters visual information processing, module grafting shows that the learned changes are distributed across both the vision and language stacks, while linear probing finds that task-relevant visual properties become more linearly accessible in middle-to-late language layers. Together, these results show that exact visual grounding can complement multimodal training by strengthening how VLMs extract and utilize fine-grained visual evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.