Credit Where It’s Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet how visual evidence is integrated during reasoning remains poorly understood. We investigate multimodal RLVR through the lens of cross-modal attention connectivity and find that only a small fraction of tokens—approximately 15%—exhibit strong visual-textual coupling. These high-connectivity tokens act as anchors that ground reasoning in the image, while most other tokens follow linguistic patterns. During RLVR training, credit assignment naturally concentrates on these anchors, progressively sharpening their visual grounding. Based on this insight, we propose Anchor-Token Reinforcement Learning (AT-RL), a lightweight framework that selectively reinforces high-connectivity tokens through graph-based clustering of attention topology. Evaluated on the Qwen2.5-VL series (3B–32B), AT-RL introduces only 1.2% computational overhead while enabling the 32B model to surpass the 72B-Instruct baseline on MathVista, achieving a score of 80.2, with consistent improvements across STEM, video, and general tasks. In contrast, training exclusively on low-connectivity tokens causes severe performance degradation, highlighting the importance of precise credit assignment to visually grounded tokens. These results suggest that reasoning quality depends more on the fidelity of cross-modal anchoring than on token quantity alone, providing an operational account of how visual evidence is selected and reinforced during multimodal RLVR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.