Where to Apply Privilege: On-Policy Self-Distillation with Textual Privilege for Multimodal LLMs
Abstract
On-policy self-distillation (OPSD) has emerged as an effective post-training paradigm for multimodal large language models (MLLMs) that enhances visual understanding without external teacher models or reward verifiers. While vanilla OPSD adopts reference answers as textual privilege, recent variants for visual understanding primarily rely on visual privilege. However, we observe that MLLMs leverage answer-supporting evidence more effectively from targeted textual descriptions than from the original images, even when both inputs are individually sufficient for deriving the answer. This suggests substantial potential for textual privilege in self-distillation. Nevertheless, naive use of textual privilege can introduce context-specific artifacts, such as “based on the description”, even when no description is available to the student. This behavior reflects a failure to disentangle genuine improvements from traces of privileged text, which we term the privileged-context attribution problem. To address this problem, we propose Role-OPD. It adopts caption-based textual privilege to enhance the model's visual understanding. Furthermore, instead of presenting the privileged caption as user-provided context, Role-OPD places it in an intermediate assistant turn, framing it as the teacher's own visual understanding. This role-aware design suppresses the transfer of context-specific artifacts while preserving the genuine improvements from privileged captions. Experiments on six fine-grained visual understanding benchmarks show that Role-OPD achieves consistent gains and outperforms existing post-training methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.