Reward-Grounded Visual Evidence: Shaping Token-Level Attention in Multimodal Reinforcement Learning
Abstract
Reinforcement learning (RL) enhances visual reasoning in multimodal large language models (MLLMs), yet existing methods rely primarily on holistic, outcome rewards. This creates a mismatch with visual grounding, which demands spatial and token specificity. We reveal that while outcome RL improves benchmark scores, target attention remains largely unaligned. Furthermore, visual sensitivity varies sharply across generation, surging on decisive tokens like coordinates while remaining negligible on auxiliary text, rendering coarse scalar rewards insufficient for critical steps. To bridge this gap, we propose Visual-credit Induced Self-attention Target Alignment (VISTA) to provide selective, token-level visual supervision within policy optimization. Rather than intervening uniformly, VISTA derives gradient-based visual credits to construct targeted attention distributions on decisive tokens, regularizing cross-modal attention through intrinsically derived targets while focusing on decisive steps. Across grounding benchmarks, VISTA lifts 3B accuracy from 68.5% to 70.3% and 7B accuracy from 69.8% to 72.1%, with gains of +4.3 and +4.8 on LISA-G. Generalizing beyond grounding, it secures a top 54.7% average on multimodal reasoning among 7B architectures, rivaling the InternVL2.5-38B model. Mechanistically, VISTA concentrates internal cross-modal attention directly onto target regions while suppressing spurious visual noise, fostering stronger alignment between policy updates and decisive visual evidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.