Functionally Grounded Latent Distillation for Visual Latent Reasoning
Abstract
Visual latent reasoning offers a promising alternative to explicit multimodal reasoning by performing intermediate computation in continuous representation space. Existing approaches commonly learn latent states by aligning them with perceptual features or privileged intermediate representations, yet representational similarity alone does not guarantee that the learned states preserve the reasoning computation they are intended to replace. We propose HCLR, a hierarchical latent distillation framework for visual latent reasoning that moves beyond representation alignment by learning latent states according to the downstream reasoning behavior they preserve. We formulate this principle as functional equivalence and progressively internalize explicit multimodal reasoning through hierarchical latent distillation. HCLR first learns compact latent substitutes for intermediate visual observations and then compresses the remaining textual and local-latent reasoning scaffold into a compact global latent trajectory. Privileged visual observations and reasoning scaffolds are used only during training and are subsequently removed through distillation, enabling the final model to perform intermediate reasoning entirely in continuous latent space using only the original image and question. Empirical results demonstrate that HCLR achieves strong performance across diverse real-world visual perception and reasoning settings, while exhibiting strong out-of-distribution generalization on challenging abstract visual reasoning tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.