Beyond Imitation: Multi-Teacher On-Policy Distillation for Multi-Modal LLMs
Abstract
On-policy distillation (OPD) has demonstrated its effectiveness in improving multimodal reasoning with multimodal large language model (MLLM) teachers. In this paper, we investigate how a text-only LLM teacher can complement MLLM supervision in multimodal OPD. We find that directly distilling from the LLM teacher can substantially degrade student performance. Although its predictions lack direct visual grounding, they reveal which reasoning continuations are favored by textual context alone. We therefore use these predictions as a reference to constrain excessive text-only alignment rather than as targets to imitate. Based on this insight, we propose Role-Aware Heterogeneous Teacher Guidance (RouTe), which combines distillation from the MLLM teacher at visually dependent positions with controlled separation from the LLM teacher's prediction distribution at high-entropy positions in reasoning trajectories. RouTe dynamically adjusts the separation strength according to the student's relative divergence from the two teachers. Experiments across multiple multimodal reasoning benchmarks and student–teacher configurations demonstrate consistent improvements over existing OPD and multimodal distillation methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.