Ask a Better Question: Unlocking What Vision-Language Models Already Get Right
Abstract
Vision-language models are queried through language, yet a frozen model's answer can vary substantially across questions intended to express the same visual request. This sensitivity creates both a practical failure mode and an opportunity: the same model may answer an alternative formulation correctly after failing on the original question. We ask whether the model's own first response can guide a better formulation at inference time. We introduce Inference-time response-CONditioned reframing (ICON), a lightweight text rewriter that observes the original question, answer options when present, and the frozen VLM's initial response, then produces a single revised question for a second pass through the same model. is trained offline from supervised rewrite pairs, but deployment requires neither gold labels nor candidate search, and the target VLM remains frozen throughout. On Robo2VLM, relative to question-only rewriting, the full response-conditioned rewriter increases the mean accuracy gain from to percentage points with a B rewriter and from to points with an B rewriter. In the matched comparison, adding the initial response to question-and-options-only rewriting contributes an additional and points, respectively. Across eight benchmarks, separately trained rewriters yield gains of to percentage points, and a Robo2VLM-trained rewriter transfers across several frozen VLMs on the same task. Together, these results identify the model's own response as a useful conditioning signal for targeted question reformulation, enabling the same frozen VLM to produce correct answers under revised formulations when the original wording leads to an error.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.