acceptodds
Under review as a conference paper at ICLR 2027

Beyond Output Alignment: Capability Decomposed On Policy Latent Distillation for Vision Language Reasoning

Abstract

On-policy distillation (OPD) has achieved strong results in LLMs by providing dense teacher supervision on student-visited states, reducing the train–inference mismatch. However, OPD supervision is mediated through next-token distributions: it teaches the student how the teacher would act, but does not explicitly transfer the intermediate representations supporting visual grounding and reasoning. Distilling the internal visual-reasoning states is promising, yet it faces two critical challenges: teacher and student representations may fall into different spaces, and their reasoning behaviors can differ substantially in wording, length, and granularity. Addressing these challenges, we introduce CD-OPLD, a capability-decomposed framework for latent distillation in VLMs. CD-OPLD starts from a structured cold-start checkpoint, then learns capability-specific visual and reasoning projectors to align student and teacher representations. During on-policy training, student-generated trajectories are paired with answer-correct teacher reference trajectories, and a monotonic semantic block matcher enables direct latent alignment between semantically corresponding visual and reasoning states. Distilling Qwen3.5-9B into Qwen3.5-2B, CD-OPLD consistently improves over a cold-start-matched OPD baseline across seven multimodal benchmarks. Code is available at https://anonymous.4open.science/r/CD-OPLD-889F/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.