acceptodds
Under review as a conference paper at ICLR 2027

LatentOPD: On-Policy Distillation for Multimodal Reasoning with Visual Latents

Abstract

Visual latent reasoning lets multimodal language models carry visual evidence through continuous tokens interleaved with textual reasoning, reducing reliance on external tools and explicit image generation. Existing training approaches often supervise these tokens with fixed latent trajectories from image–text reasoning traces. However, uninformative auxiliary views can encourage shared soft prompts, while fixed targets fail to adapt to the student's evolving reasoning context. We propose , which combines informative visual supervision with on-policy distillation over latent sub-trajectories under teacher-forced textual contexts. We train a privileged teacher to encode information-rich auxiliary views into latent trajectories that support subsequent reasoning. After learning to reason from teacher-provided latents, the student learns to generate latent states supervised by the teacher's privileged representations conditioned on student-generated latent prefixes. At inference, the student alternates latent and textual reasoning using only the original query. Experiments with Qwen2.5-VL-7B show improved accuracy over supervised fine-tuning baseline across four benchmarks of visual, mathematical, and geometric reasoning, with a higher weighted average than the compared visual latent methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.