acceptodds
Under review as a conference paper at ICLR 2027

PReViD: Interventional Prefix Replay for Visual On-Policy Distillation

Abstract

Visual distillation uses a teacher with better visual evidence to supervise responses sampled by a student model, but early visual mistakes remain in the response history and distort later supervision. We identify this missing pathway by showing that corrected visual evidence can change later predictions through the teacher's language context. Based on this finding, we introduce PReViD, which replays the same student response under controlled image and evidence contexts while preserving alignment with the original response. PReViD separates visual and contextual corrections, uses a placebo prefix with no visual evidence to filter generic text effects, and gives more weight to context changes that support the teacher's correction. Across benchmarks for detailed visual understanding, model sizes, and teacher settings, PReViD outperforms strong distillation baselines without changing inference. Code is available at https://anonymous.4open.science/r/E02D2.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.