acceptodds
Under review as a conference paper at ICLR 2027

Seeing Before Distilling: Prefill Visual Alignment for Multimodal On-Policy Distillation

Abstract

On-policy distillation (OPD) mitigates the distribution mismatch of offline distillation by training students on their own generated trajectories. However, we find that vanilla OPD is often insufficient for vision-language models, especially on visually dominated tasks, where students may attend to irrelevant image regions and weaken subsequent token-level distillation. To address this issue, we propose Prefill Visual Alignment On-Policy Distillation, which aligns the student’s visual-token attention with the teacher during the prefill stage. This avoids task-specific decoded-token selection and provides a simple, unified supervision signal across multimodal tasks. Experiments on InternVL3 and Qwen3-VL show that our method improves on most benchmarks over vanilla OPD, e.g., from 44.8 to 51.5 on InfoVQA and from 69.1 to 71.8 on OCRBench. These results demonstrate that prefill-stage visual alignment strengthens visual grounding and enables more effective multimodal knowledge transfer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.