acceptodds
Under review as a conference paper at ICLR 2027

Self-Verifying Vision Language Models

Abstract

Reinforcement learning with verifiable rewards has emerged as a powerful paradigm for improving reasoning in structured domains such as mathematics and code, where deterministic checkers can directly certify generated solutions. Extending this paradigm to visual reasoning is difficult because correctness requires interpreting unstructured visual evidence, and ground-truth is scarce at scale. We introduce a training framework for vision-language models (VLMs) that makes verification a learned capability rather than relying on exact rewards to verify final answers. The model generates a candidate solution, critiques its own reasoning, and then refines its answer, while a simple proxy reward evaluates the resulting generation–verification pair. We find that jointly training for generation and verification creates a positive feedback loop during training that yields more scalable visual reasoning than optimizing answer correctness alone, even when self-verification is disabled at test time. Across diverse benchmarks, our framework improves the base VLM, outperforms answer-supervised training, and surpasses test-time scaling baselines, with the largest gains on 3D-grounded reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.