P-GVAR: Selective Visual Re-Grounding under Image-Text Grounding Drift
Abstract
Vision–language models (VLMs) can follow misleading textual context even when the image supports a different answer. We study a simple recovery mechanism: querying the same frozen VLM again after removing potentially conflicting context. This fresh visual query repairs many context-induced errors, but unconditional context removal is not generally safe because reliable context can also improve prediction. We therefore formulate recovery as a selective visual re-grounding problem with two distinct components: constructing a fresh visual candidate and deciding whether it should replace the original response. We introduce Policy-Guided Visual Answer Refresh (P-GVAR), a training-free response-level policy that uses controlled cross-condition responses to make this retain-or-refresh decision without task labels or access to model internals. Across three open VLMs and four stress benchmarks, selective re-grounding preserves most of the gain of unconditional visual reset, while controlled context experiments show that arbitration becomes important when context can be either helpful or misleading. These findings motivate evaluating candidate quality and routing quality separately when studying recovery from image–text grounding drift.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.