acceptodds
Under review as a conference paper at ICLR 2027

P-GVAR: Selective Visual Re-Grounding under Image-Text Grounding Drift

Abstract

Vision–language models (VLMs) can follow misleading textual context even when the image supports a different answer. We study a simple recovery mechanism: querying the same frozen VLM again after removing potentially conflicting context. This fresh visual query repairs many context-induced errors, but unconditional context removal is not generally safe because reliable context can also improve prediction. We therefore formulate recovery as a selective visual re-grounding problem with two distinct components: constructing a fresh visual candidate and deciding whether it should replace the original response. We introduce Policy-Guided Visual Answer Refresh (P-GVAR), a training-free response-level policy that uses controlled cross-condition responses to make this retain-or-refresh decision without task labels or access to model internals. Across three open VLMs and four stress benchmarks, selective re-grounding preserves most of the gain of unconditional visual reset, while controlled context experiments show that arbitration becomes important when context can be either helpful or misleading. These findings motivate evaluating candidate quality and routing quality separately when studying recovery from image–text grounding drift.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.