Seeing while Reasoning: Advancing Visual Faithfulness in Small Multi-modal Language Models with Verifiable Perception-aware Intermediate Process Optimization
Abstract
Reinforcement learning with verifiable rewards (RLVR) advances multimodal chain-of-thought (MCoT) reasoning, but prevailing objectives emphasize answer correctness or output-level vision-language consistency, leaving intermediate reasoning weakly supervised for visual faithfulness. This work analyze direct text-to-visual attention using Visual Attention Mass (VAM) and identify two systematic behaviors. Visual Attenuation denotes the abrupt reduction in visual attention when generation transitions from captioning to reasoning. Perception Sensitivity captures the positive association between stage-level VAM and the correctness of the corresponding intermediate stage. Motivated by these findings, we operationalize the Seeing while Reasoning principle and construct MCoT-ViG, a structured MCoT dataset pairing four-stage rationales with controlled visual variants and object-level grounding annotations. We further introduce Perception-aware Intermediate Process Optimization (PaIPO), which combines stage-wise process rewards for intermediate supervision with two complementary perception rewards. Caption Perception Reward (CPR) measures caption sensitivity to controlled visual variants, whereas Reasoning Grounding Reward (RGR) verifies object references against annotated target regions. Experiments across five benchmarks and three Qwen3-VL model scales demonstrate that PaIPO consistently improves multimodal reasoning and visual faithfulness. Relative to the backbones, PaIPO achieves absolute improvements of +3.53 to +6.14 in VRC-Bench Step Score and attains edited-image accuracies of 55.92% to 65.69% on VFaith-Bench. Ablations confirm the complementary contributions of CPR and RGR, while analysis of reward dynamics and post-training VAM demonstrate faster convergence and more sustained visual attention throughout intermediate reasoning, respectively. Code and dataset will be available to facilitate reproducible research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.