acceptodds
Under review as a conference paper at ICLR 2027

Structured Latent Reasoning: Refactoring Reasoning with Reusable Primitives for Multimodal LLMs

Abstract

Chain-of-Thought reasoning has significantly advanced problem solving in multimodal large language models (MLLMs), yet autoregressive generation of lengthy reasoning traces incurs substantial token-level redundancy and computational overhead. Existing approaches compress reasoning into continuous latent states or discrete latent representations, but typically represent heterogeneous reasoning content within a shared latent space, making it difficult to efficiently encode both reusable reasoning patterns and instance-specific information. In this paper, we present **SLR**, a **S**tructured **L**atent **R**easoning framework that separates *reusable reasoning operations* from their *instance-specific executions*. SLR segments reasoning traces into unit spans, decomposes each span into an abstract operation and its concrete execution, and clusters recurring operations to construct a compact discrete *Reasoning Codebook*. During training, operation descriptions are replaced by discrete codebook tokens and interleaved with explicit execution text, allowing recurring procedural patterns to be represented compactly while preserving problem-specific computations. We evaluate SLR across diverse multimodal reasoning benchmarks with three MLLM configurations. Compared with standard SFT, SLR reduces mean output length by 10-35%, while matching or exceeding standard SFT accuracy in all settings. These results demonstrate that explicitly separating reusable reasoning structure from instance-specific execution provides an effective approach to reducing reasoning redundancy while maintaining multimodal reasoning performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.