Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
Abstract
Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and recent work extends RLVR to multimodal large language models (MLLMs). This transfer, however, surfaces a faithfulness challenge: faithful perception of task-relevant visual evidence and faithful use of that evidence during reasoning. Existing perception supervision often operates on textual descriptions rather than natively on image regions, while answer-level rewards provide limited guidance on when to attend to visual evidence during reasoning. We introduce Faithful-MR1, an anchor-then-reinforce training framework that guides where and when to attend to visual evidence through question-relevant spatial supervision and intervention-guided attention reinforcement. The Anchoring stage turns perception into an explicit pre-reasoning subtask, supervising a dedicated <Focus> token's attention directly against image regions rather than through textual descriptions. The Reinforcing stage uses counterfactual image intervention to identify vision-dependent response tokens, rewarding answer-correct trajectories that concentrate visual attention at those positions. Extensive experiments demonstrate that Faithful-MR1 outperforms recent multimodal reasoning baselines on both Qwen2.5-VL-Instruct 3B and 7B backbones while using substantially less training data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.