VC-RL: Visual Contrast-Guided Reinforcement Learning for Native Visual Thinking
Abstract
Fine-grained visual understanding requires multimodal large language models to identify local details accurately and use them in visual reasoning. Native thinking can improve this capability, but the use of local visual feedback to improve long-form reasoning over full images remains underexplored. Local views improve a teacher's access to fine details while also changing image scale and context, so differences between teacher and student predictions can reflect both target evidence and other changes in visual conditions. We propose Visual Contrast-Guided Reinforcement Learning (VC-RL), which combines Group Relative Policy Optimization (GRPO) with visual on-policy self-distillation. The student generates thinking trajectories from full images, and a self-teacher scores the same trajectories under size-matched positive-evidence and counterfactual views constructed with shared preprocessing rules. The contrast between these predictions guides token selection and bounded modulation of group-relative advantages derived from sequence-level rewards, without reversing the signs of nonzero advantages. Regional visual feedback thereby contributes selectively to token-level optimization, while task-level rewards constrain the policy updates. Across six fine-grained and high-resolution visual evaluations, VC-RL achieves average accuracies of 79.21% and 81.33% with Qwen3.5-4B and Qwen3.5-9B, respectively, improving over the corresponding base models by 4.45 and 3.90 percentage points, with additional gains on multiple general multimodal benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.