V-ERR: Visual Evidence Re-checking Reinforcement in Multimodal Reasoning
Abstract
Despite their strong perception and generation abilities, vision-language models (VLMs) still struggle with complex reasoning. While GRPO-style post-training reinforcement learning has improved multimodal reasoning performance, it primarily optimizes external outcomes and underutilizes the model’s internal modality signals, especially the visual signals that should be tracked and verified across reasoning steps. To address these challenges, we draw inspiration from cognitive science on multimodal signal processing, and uncover two key phenomena in VLMs: signal partitioning, where modality signals are processed through partitioned pathways, and signal attenuation, where modality-specific signals progressively decays as reasoning unfolds. Building on these insights, we propose V-ERR (**V**isual **E**vidence **R**e-checking **R**einforcement), a post-training framework that reinforces attenuated signals during reasoning. Concretely, we first introduce a partitioning score to identify modality-specialized attention pathways and then incorporate GRPO with modality-specialized reward shaping to reinforce evidence-checking to strengthen the attenuated visual-modality signals. Extensive experiments demonstrate that V-ERR significantly enhances the reasoning capabilities of VLMs, consistently outperforming leading RL optimization strategies by a substantial margin. The code is available at https://anonymous.4open.science/r/V-ERR-BB96.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.