VC-OPD: Learning to See What Matters via Visual Cue On-Policy Self-Distillation
Abstract
Existing visual On-Policy Distillation (OPD) methods typically use cropped or zoomed-in image regions as teacher-side privileged information to improve the fine-grained visual perception of small Multimodal Large Language Models (MLLMs). However, such local pixel-level information is insufficient for complex visual reasoning that requires understanding global structures, spatial relationships, and high-level semantics. To address this limitation, we propose Visual Cue On-Policy Self-Distillation (VC-OPD), which extends teacher-side privileged information from local image regions to question-conditioned visual cues. Specifically, we leverage large advanced MLLMs to extract the key visual evidence required to solve a given question, including key visual details, structural relationships, and high-level semantics. A multi-stage quality filtering mechanism is further introduced to retain reliable visual cues as privileged information. During training, these cues are provided only to the teacher model, enabling it to deliver more informative token-level supervision along student-generated trajectories. The student is thereby encouraged to internalize the ability to identify and reason over critical visual evidence directly from the original image, without requiring any additional information at inference time. Since large advanced MLLMs are used only for visual cue generation, VC-OPD is applicable to both open-weight and proprietary black-box models. Experiments on STEM and general visual reasoning benchmarks show that VC-OPD consistently improves visual reasoning performance across different model scales and outperforms a range of reinforcement learning and on-policy distillation baselines. The 4B model trained with VC-OPD surpasses many baselines with 8B or more parameters, demonstrating the effectiveness of question-conditioned visual cues as privileged information for capability transfer. The code will be released after review.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.