acceptodds
Under review as a conference paper at ICLR 2027

Visual Contrastive Self-Distillation

Abstract

On-policy self-distillation (OPSD) removes the external teacher required by on-policy distillation (OPD), but still needs asymmetric information between teacher and student so that the self-teacher provides a stronger learning signal. Existing methods create this asymmetry through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. We propose Visual Contrastive Self-Distillation, namely VCSD, which converts image-content removal into an on-policy self-distillation signal. At each student-generated response prefix, the EMA teacher produces two next-token distributions under the same prompt and prefix: one conditioned on the original image and the other on a content-erased control. Their token-wise log-probability difference highlights candidates whose likelihood is specifically increased by the instance-level visual content. We use this contrast to sharpen the teacher's original-image distribution within its plausible support, and distill the resulting full-distribution target into the student. On ViRL39K, VCSD consistently outperforms matched OPSD across Qwen3-VL and Qwen3.5 models. It improves the seven-benchmark aggregate from 72.51% to 76.26% on Qwen3-VL-8B and from 74.97% to 79.24% on Qwen3.5-9B, with gains at all six evaluated scales. VCSD requires no external teacher, privileged answers, visual evidence signals, reasoning traces, or additional inference-time cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.