acceptodds
Under review as a conference paper at ICLR 2027

Self-Verifiable Multimodal Reasoning via Adversarial Co-Evolution

Abstract

Reliable multimodal reasoning requires answers that are supported by the model's own reasoning. A strong external judge can supervise this consistency, but evaluating every rollout with a large model is costly. A compact policy that verifies its own output offers an attractive alternative, yet its judgments require a reliable learning signal. We propose **Self-Verifiable Policy Optimization (SVO)**, which transfers external verification supervision into a co-evolving discriminator and feeds its learned score back into policy training. The policy generates a reasoning trace, an answer, and a self-verification judgment. Periodically, a frozen judge supplies verification references for a small subset of current rollouts; the discriminator learns to prefer these references and provides both a dense reward and a gate on answer-correctness rewards. This couples learning to assess consistency with learning to generate consistent responses, while reducing repeated large-judge calls. Across four Qwen2.5-VL and Qwen3-VL backbones and seven multimodal math benchmarks, SVO improves reasoning–answer consistency and consistency-constrained accuracy over GRPO by **14.7 and 6.2 percentage points** on average. It approaches direct online judge supervision, with gaps of 1.1 and 0.6 points, respectively. Comparisons with consistency-aware baselines and additional reasoning tasks support SVO as a practical approach to learning from external verification with fewer judge queries.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.