Weak-to-Strong On-Policy Distillation via Perception Gap Supervision
Abstract
On-policy distillation (OPD) supervises a student on its own rollouts using token-level targets from a teacher. Existing OPD methods typically rely on a teacher model that is stronger than the student. Training such a teacher can be expensive or even impractical. We study weak-to-strong OPD for vision-language models, where the teacher is weaker than the student, typically being smaller or less extensively trained. Directly distilling the weak teacher's predictions can degrade student performance. We introduce Weak-to-Strong Distillation via Perception Gap (W2S-PG), which constructs supervision from the discrepancy between the same weak teacher's predictions under two answer-preserving visual perturbations. While teacher-derived supervision is effective, the student becomes stronger during training, and its own predictions may eventually provide more reliable supervision than the teacher-derived signal. To select the more reliable source of supervision, we further propose Adaptive Supervision Signal Selection (AS3), a token-level mechanism that adaptively chooses between teacher-derived perception-gap supervision and self-training according to model confidence. Across six perception benchmarks and four teacher–student pairs, W2S-PG consistently improves the student and outperforms other OPD variants such as vanilla OPD and previous weak-to-strong distillation approaches. When combined with W2S-PG, AS3 provides further gains of 0.3–2.4 points across all benchmarks. Ablations of augmentation and discrepancy choices demonstrate the robustness of W2S-PG, and evaluations on seven additional benchmarks show that its benefits extend beyond perception to visual reasoning and other tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.