Seeing Before Distilling: Prefill Visual Alignment for Multimodal On-Policy Distillation
Abstract
On-policy distillation (OPD) mitigates the distribution mismatch of offline distillation by training students on their own generated trajectories. However, we find that vanilla OPD is often insufficient for vision-language models, especially on visually dominated tasks, where students may attend to irrelevant image regions and weaken subsequent token-level distillation. To address this issue, we propose Prefill Visual Alignment On-Policy Distillation, which aligns the student’s visual-token attention with the teacher during the prefill stage. This avoids task-specific decoded-token selection and provides a simple, unified supervision signal across multimodal tasks. Experiments on InternVL3 and Qwen3-VL show that our method improves on most benchmarks over vanilla OPD, e.g., from 44.8 to 51.5 on InfoVQA and from 69.1 to 71.8 on OCRBench. These results demonstrate that prefill-stage visual alignment strengthens visual grounding and enables more effective multimodal knowledge transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.