CV-OPD: Contrastive Vision On-policy Distillation
Abstract
On-policy distillation (OPD)reveals versatility and convenience during VLM post training, providing dense supervision for student from teacher. However, we ob served an anomaly that the distilled VLM could still answer the questions accu rately even when the related visual regions were masked or replaced with noise . Building on this observation, our quantitative and qualitative experiments demon strate that OPD, which uses the answer as a hint, has failed to teach students to genuinely seek out and rely on visual evidence . Our theoretical analysis suggests that this failure arises from a shortcut enabled by the lack of visual constraints dur ing distillation . To address the issue, we introduce Contrastive Vision On-policy Distillation (CV-OPD), which compares teacher distributions induced by positive and negative crops while conditioning on the same student prefix and augments optimization with a margin-based contrastive constraint, which is modulated with discrepancy gating across token and sample level . Specifically, CV-OPD first introduces a negative teacher unrelated to the answer and designs a Contrastive Divergence Loss (CDL) that constrains the feature distribution of student by con trasting the positive and negative teachers, correcting how the student exploits visual regions . The CV-OPD introduces a Discrepancy Aware Gate (DAG) that uses token-level and sample-level discrepancies to select visually relevant tokens for optimization,directing the student’s attention toward visual content . Extensive experiments across a broad range of fine-grained visual benchmarks demonstrate that CV-OPD substantially alleviates shortcut and consistently improves the visual reasoning capabilities of existing VLMs . All the reproduction code will be re leased on https://anonymous.4open.science/r/CV-OPD-9D04/ .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.