acceptodds
Under review as a conference paper at ICLR 2027

Can Visual Chain-of-Thought Trust Its Visual Evidence? Pre-Consumption Evidence Verification for Interleaved VLMs

Abstract

Interleaved vision-language models revisit images during reasoning, yet typically consume newly selected evidence without checking whether it matches the current visual request. We show through controlled evidence interventions that such mismatches are consequential: replacing erroneous evidence with ground-truth evidence substantially improves answer accuracy, whereas counterfactual evidence degrades it. Moreover, correcting errors after consumption poses a dilemma: appending correct evidence only partially recovers accuracy, while rollback requires discarding already-generated reasoning. These findings motivate a simple principle: verify visual evidence before consumption. We introduce VERIFYCOT, which checks whether selected evidence matches the model’s current visual request. This request is derived directly from the model’s hidden states, eliminating the need to extract a textual reference. Evidence that passes is kept, and evidence that fails is corrected before reasoning continues. Across interleaved VLM architectures and five visual reasoning benchmarks, VERIFYCOT improves end-to-end accuracy in 14 of 15 evaluated settings, with gains of up to 15.7 percentage points. Compared with unconditional grounding, it achieves higher accuracy with fewer grounding calls, while also offering a better accuracy–compute trade-off than the evaluated post-hoc correction methods. We provide code in the supplementary material and will release it on GitHub.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.