acceptodds
Under review as a conference paper at ICLR 2027

Functionally Grounded Latent Distillation for Visual Latent Reasoning

Abstract

Visual latent reasoning offers a promising alternative to explicit multimodal reasoning by performing intermediate computation in continuous representation space. Existing approaches commonly learn latent states by aligning them with perceptual features or privileged intermediate representations, yet representational similarity alone does not guarantee that the learned states preserve the reasoning computation they are intended to replace. We propose HCLR, a hierarchical latent distillation framework for visual latent reasoning that moves beyond representation alignment by learning latent states according to the downstream reasoning behavior they preserve. We formulate this principle as functional equivalence and progressively internalize explicit multimodal reasoning through hierarchical latent distillation. HCLR first learns compact latent substitutes for intermediate visual observations and then compresses the remaining textual and local-latent reasoning scaffold into a compact global latent trajectory. Privileged visual observations and reasoning scaffolds are used only during training and are subsequently removed through distillation, enabling the final model to perform intermediate reasoning entirely in continuous latent space using only the original image and question. Empirical results demonstrate that HCLR achieves strong performance across diverse real-world visual perception and reasoning settings, while exhibiting strong out-of-distribution generalization on challenging abstract visual reasoning tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.