acceptodds
Under review as a conference paper at ICLR 2027

VISPO: Verification-State-Aware Policy Optimization for Reliable Multimodal Reasoning

Abstract

Reinforcement learning with verifiable rewards is crucial for enhancing the complex reasoning capabilities of multimodal large language models. However, existing paradigms primarily rely on outcome supervision, providing sparse guidance for intermediate reasoning and lacking reliable verification signals to mitigate hallucination propagation during iterative reasoning. To address these limitations, we propose VISPO, Verification-State-Aware Policy Optimization, which trains an agentic verifier to acquire evidence through tool interactions, perform evidence-grounded verification across support, insufficient, and contradict states, and generate reliable verification explanations for multimodal reasoning. Moreover, VISPO introduces region-level advantage redistribution to enable state-aware credit assignment and alleviate credit misalignment. To provide structured supervision for verifier training, we construct FGMV, a Fine-Grained Multimodal Verification dataset with process-level state-aware verification annotations and explanations. Building upon VISPO, we further propose VISPO-R, a synergistic multimodal reasoning framework that leverages iterative verifier feedback via verification-aware hierarchical advantage assignment for reasoning refinement. Extensive experiments show that VISPO-R consistently improves Qwen3-VL backbones across both model scales, achieving average relative improvements of 13.6%, 5.5%, and 10.3% on mathematical reasoning, general reasoning, and hallucination benchmarks. These results demonstrate that state-aware verification effectively improves the accuracy and reliability of multimodal reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.