Do Latent States Really Reason? Making It Great Again through Information Control
Abstract
Textual chain-of-thought (CoT) improves multimodal reasoning but incurs substantial autoregressive decoding cost. Latent reasoning reduces this overhead by compressing intermediate reasoning into a fixed number of latent states. Despite recent progress, existing methods often implicitly assume that supervised latent states necessarily encode useful reasoning information. Our empirical results challenge this assumption, showing that latent states do not necessarily encode meaningful semantic information and that even skipping them during inference can have little impact on performance. These findings call into question whether current latent reasoning methods truly learn informative and interpretable latent representations. To address this, we propose ICLR (**I**nformation **C**ontrol for **L**atent **R**easoning), which explicitly controls information flow to make latent states both necessary for prediction and effective for reasoning. Specifically, we design a two-stage framework. In the first stage, we explicitly route all information through the latent states, making them necessary intermediates in the reasoning pathway. In the second stage, we further extract the visual evidence and textual semantics required for step-level reasoning and reinforce the latent states to encode this information, thereby enabling effective latent reasoning. Extensive experiments show that ICLR achieves state-of-the-art performance across multimodal benchmarks, surpassing existing latent reasoning baselines and latent visual reasoning approaches requiring auxiliary images and costly annotations. Compared to textual CoT methods, ICLR achieves an average performance gain of **3.6%** while delivering a **20.9×** generation speedup. Code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.