acceptodds
Under review as a conference paper at ICLR 2027

Right Answers for the Right Reasons? Bridging Reasoning Utility and Visual Dependency in Multimodal Reinforcement Learning

Abstract

Reinforcement learning with verifiable rewards (RLVR) has advanced multimodal reasoning, but final-answer supervision cannot distinguish sound reasoning from flawed trajectories that happen to produce correct answers. Our model-based and human assessments reveal substantial perception and reasoning errors among answer-correct responses, motivating feedback that differentiates successful trajectories. We propose Centered-Adaptive Residual Visual Transport (CA-RVT), a design that evaluates reasoning traces using the policy's own answer-confidence probes. CA-RVT measures how each trace changes reference-answer confidence with and without direct image access, relative to trace-free baselines. It combines multimodal reasoning utility with a conservative cross-view residual and adapts the latter to visual dependency across training problems. By reusing the policy for confidence scoring, it requires neither process annotations nor external reward or judge models during training, keeping the post-training pipeline lightweight. With Qwen2.5-VL-7B-Instruct post-trained on only 2.1K Geometry3K problems, CA-RVT achieves competitive performance across mathematical reasoning and general multimodal benchmarks. It attains the highest four-benchmark mathematical average among the compared open-source models with complete results, and the best results on all seven general multimodal benchmarks in our comparison. On questions answered correctly by all audited models, it also achieves lower judged spurious-correctness rates. Additional experiments across training datasets and different backbones further support its applicability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.