Visual Evidence Prototype-Guided Policy Optimization for Visual Reasoning
Abstract
Reinforcement learning with verifiable rewards has improved visual reasoning in large vision-language models. Existing methods either rely on visual dependence, while strong visual dependence does not imply correct use of visual information, or require external supervision over task-relevant image regions. To address these issues, we introduce Visual Evidence Prototypes (VEPs), which derive fine-grained visual supervision from on-policy rollout outcomes without region annotations. Specifically, for each question, VEPs aggregate the attention distributions of correct and incorrect rollouts over visual tokens into Success and Failure prototypes, and contrast them to capture the key differences. Empirical analyses show that VEPs identify key visual and response tokens for visual reasoning. Intervening on only the top 10% visual tokens selected by VEPs reduces answer accuracy by 20.15–28.45 percentage points. Based on these findings, we propose VEP-RL, which uses VEP to reweight trajectory-level updates and redistribute token-level credit to selected tokens. Across fourteen visual reasoning and perception benchmarks, VEP-RL improves average accuracy over the strongest baseline for each backbone by 2.44, 1.93, and 3.24 absolute percentage points on three models, respectively, with its effectiveness consistently validated in out-of-distribution benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.