acceptodds
Under review as a conference paper at ICLR 2027

Premise over Pixels: Diagnosing Visual Self-Correction in Vision-Language Models

Abstract

Multi-step reasoning requires a model to overturn its own earlier errors before they propagate. In vision-language models (VLMs), the same image remains available even after the model has stated an incorrect visual premise. We introduce false-premise replay to test whether VLMs can use that image to revise their own errors. From a failed reasoning trace, we extract one answer-relevant visual claim contradicted by the image, discard the remaining reasoning and final answer, and replay the claim with the original image and question. Across three model-reviewed evaluation sets, replay lowers accuracy relative to a fresh query by 25.8–31.0 percentage points, while substituting a corrected premise raises it. Image interventions show continued visual influence, but models do not reliably identify the image-supported premise. On InternVL3.5-8B, a false premise also reduces how often sampled answers include a correct one and impairs selection from fixed candidate pools. We consolidate 217 premise pairs into a shared evaluation set and compare correction baselines on five distinct models. On this set, false-premise replay lowers accuracy by 16.6 points even for a 975B-parameter model with reasoning enabled. On all five models, same-model visual-verification feedback improves accuracy over self-reflection with false premises, but lowers it when the earlier premise is correct. The findings distinguish obtaining corrective information from using it, and show why measuring repair under false-premise replay alone can overstate correction quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.