acceptodds
Under review as a conference paper at ICLR 2027

Credit Where It’s Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet how visual evidence is integrated during reasoning remains poorly understood. We investigate multimodal RLVR through the lens of cross-modal attention connectivity and find that only a small fraction of tokens—approximately 15%—exhibit strong visual-textual coupling. These high-connectivity tokens act as anchors that ground reasoning in the image, while most other tokens follow linguistic patterns. During RLVR training, credit assignment naturally concentrates on these anchors, progressively sharpening their visual grounding. Based on this insight, we propose Anchor-Token Reinforcement Learning (AT-RL), a lightweight framework that selectively reinforces high-connectivity tokens through graph-based clustering of attention topology. Evaluated on the Qwen2.5-VL series (3B–32B), AT-RL introduces only 1.2% computational overhead while enabling the 32B model to surpass the 72B-Instruct baseline on MathVista, achieving a score of 80.2, with consistent improvements across STEM, video, and general tasks. In contrast, training exclusively on low-connectivity tokens causes severe performance degradation, highlighting the importance of precise credit assignment to visually grounded tokens. These results suggest that reasoning quality depends more on the fidelity of cross-modal anchoring than on token quantity alone, providing an operational account of how visual evidence is selected and reinforced during multimodal RLVR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.