acceptodds
Under review as a conference paper at ICLR 2027

VC-RL: Visual Contrast-Guided Reinforcement Learning for Native Visual Thinking

Abstract

Fine-grained visual understanding requires multimodal large language models to identify local details accurately and use them in visual reasoning. Native thinking can improve this capability, but the use of local visual feedback to improve long-form reasoning over full images remains underexplored. Local views improve a teacher's access to fine details while also changing image scale and context, so differences between teacher and student predictions can reflect both target evidence and other changes in visual conditions. We propose Visual Contrast-Guided Reinforcement Learning (VC-RL), which combines Group Relative Policy Optimization (GRPO) with visual on-policy self-distillation. The student generates thinking trajectories from full images, and a self-teacher scores the same trajectories under size-matched positive-evidence and counterfactual views constructed with shared preprocessing rules. The contrast between these predictions guides token selection and bounded modulation of group-relative advantages derived from sequence-level rewards, without reversing the signs of nonzero advantages. Regional visual feedback thereby contributes selectively to token-level optimization, while task-level rewards constrain the policy updates. Across six fine-grained and high-resolution visual evaluations, VC-RL achieves average accuracies of 79.21% and 81.33% with Qwen3.5-4B and Qwen3.5-9B, respectively, improving over the corresponding base models by 4.45 and 3.90 percentage points, with additional gains on multiple general multimodal benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.