acceptodds
Under review as a conference paper at ICLR 2027

VERA: Visual Evidence Reinforcement for Multimodal Reasoning

Abstract

Multimodal reasoning large language models (MRLMs) have made substantial progress, yet their reinforcement learning objectives still rely primarily on final-answer correctness, leaving intermediate visual evidence largely unsupervised. Such outcome-level feedback cannot determine whether a response is grounded in the image, preserves the information needed to solve the task, or reaches a correct answer through unreliable evidence. We introduce VERA, a visual evidence reinforcement framework that externalizes task-oriented visual evidence as a structured description separated from the reasoning process and final answer. VERA evaluates this evidence along two complementary dimensions: visual faithfulness, which requires every visual claim to be supported by the input image, and task sufficiency, which requires the question and evidence alone to recover the target answer. The two evidence rewards are combined with answer correctness and format compliance in group-relative policy optimization, allowing trajectories with similar answer outcomes to receive different credit according to their visual-faithfulness and task-sufficiency assessments. Across five MRLM backbones and eight multimodal reasoning benchmarks, VERA improves average accuracy by percentage points while reducing output-token consumption by –.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.