acceptodds
Under review as a conference paper at ICLR 2027

MagistrateOPD: Verify Before Distill for Reliable On-Policy Distillation

Abstract

Existing On-Policy Distillation (OPD) methods combine policy learning from student-generated trajectories with dense token-level teacher guidance. Subsequent methods further strengthen this supervision by equipping a positive teacher with an evidence-centered zoom-in crop and constructing a negative teacher from a masked image. However, such dense supervision implicitly relies on the teacher responses being sufficiently reliable. For smaller models such as 4B-scale multi-modal large language models(MLLMs), additional zoom-in evidence may improve the positive teacher's correct-response rate while its absolute accuracy remains relatively low, substantially weakening the effective supervision signal available to the student. In this paper, we propose MagistrateOPD, which applies pre-judge verification to both teacher responses before token-level on-policy distillation. The magistrate uses an MLLM to compare both teacher responses with the reference answer at the semantic level before admitting a sample into distillation. Pre-judge verification requires the positive-teacher response to match the reference answer and the negative-teacher response not to match it. In addition, we introduce a negative-teacher loss that pushes the student away from the masked-image negative-teacher distribution, thereby strengthening the negative-teacher supervision signal. We evaluate MagistrateOPD on V* Bench, ZoomBench, HR-Bench, and MME-RealWorld. MagistrateOPD achieves 92.15 on V* Bench, 84.12 on HR-Bench-4K, and 72.94 on MME-RealWorld-EN, as well as 61.89 on ZoomBench. Across all six evaluation splits, MagistrateOPD achieves an average score of 78.02, outperforming the base model, SFT, and OPSD by 13.72, 11.17, and 8.76 points, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.