acceptodds
Under review as a conference paper at ICLR 2027

Guiding Selective Self-Correctionin Vision-Language Modelsthrough Evidence Auditing

Abstract

Vision-language models can produce coherent reasoning, but misread visual information may persist through later reasoning and self-checking. Revision can also turn correct answers into wrong ones. We propose PACT, a training-free framework with context-isolated evidence auditing and selective revision. Perception, option-level auditing, and task-adapted consistency testing run without the initial answer or its reasoning. Fusion hides earlier answer identities and combines the image with diagnostic reports. A deterministic gate accepts or rejects its candidate based on consistency and contradiction. All stages share one frozen model, without a separately trained verifier or reference answers at inference time. Across 19 open-source model configurations on MMStar and RealWorldQA, PACT scores higher than six ablations in all 38 model–dataset pairs. Its mean differences from the listed Baseline reference scores are 3.09 and 3.12 percentage points, respectively. On both benchmarks, perception has the highest single-module mean score, while removing auditing causes the largest mean loss. These results show that a module's effect depends on the full configuration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.