Distilling One Token is Enough: Rethinking On-Policy Distillation for Visual Perception in MLLMs
Abstract
On-policy distillation (OPD) has recently emerged as an effective approach for improving visual perception in multimodal large language models (MLLMs) by aligning teacher and student predictive distributions over student-generated trajectories. In visual perception tasks, these trajectories typically contain descriptive tokens that verbalize visual evidence and decision tokens that produce the final answer. This raises a fundamental question: is dense token-level supervision across the entire trajectory necessary? In this work, we find substantial redundancy in OPD supervision for visual perception. Distilling only a single decision token or only the descriptive tokens on multiple-choice visual perception data achieves downstream performance comparable to full-response distillation. Moreover, replacing the teacher distribution at the decision position with a hard ground-truth label fails to reproduce the improvement, highlighting the importance of the teacher's soft predictive distribution. Motivated by these findings, we propose **D**ecision-**O**nly **D**istillation (**DOD**), a simple, rollout-free distillation strategy that uses a direct-answer prompt and matches teacher and student distributions only at the first output position. We further introduce target-vocabulary distillation, which restricts distribution matching to the candidate decision tokens (e.g., A, B, C, and D) rather than the full vocabulary. Consequently, each training example requires distillation at only one token position over only a handful of vocabulary entries, eliminating rollout and reducing distillation costs while achieving better overall visual perception performance than standard OPD. Code is available at the [anonymous link](https://anonymous.4open.science/r/dod_code-6337).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.