acceptodds
Under review as a conference paper at ICLR 2027

Restore Text Without Breaking Vision: Vision-Preserving On-Policy Distillation for VLMs

Abstract

Vision-language models (VLMs) are built by augmenting a pre-trained large language model (LLM) with a vision encoder and fine-tuning on multimodal data, but this process often weakens the language capabilities inherited from the base LLM. A natural remedy is to distill the base LLM back into the VLM. On-policy distillation (OPD), which trains the student on its own sampled responses under teacher supervision, appears well suited to this setting because the VLM student and its base LLM teacher already have closely aligned output distributions. However, this intuition is incomplete in multimodal models. When applied with text-only rollouts, OPD improves language performance but substantially degrades vision performance, revealing a failure mode we term cross-modal distillation interference. We attribute this to a shared-parameter asymmetry. In text-only OPD, the student generates rollouts in pure-text mode, which implicitly regularizes its text-mode policy and keeps it close to the original student. In contrast, the vision-conditioned policy is not directly sampled during text-only OPD, but still shares the parameters updated by text-mode rollouts. It therefore absorbs these updates without direct regularization, leading to substantially greater drift than the text-mode policy. Motivated by this analysis, we propose Vision-Preserving On-Policy Distillation (VP-OPD), a structure-aware and mechanism-driven approximation to constrained distillation that restricts text-only updates to parameters less critical for vision-conditioned behavior. VP-OPD estimates structured vision importance on a small multimodal calibration set, freezes high-importance components, and performs distillation only on the remaining subnetwork. Experiments show that, with only 40% trainable parameters, VP-OPD recovers 95% of the language gains of full-parameter OPD while substantially preserving vision performance, achieving a significantly better text–vision tradeoff.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.