acceptodds
Under review as a conference paper at ICLR 2027

Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains

Abstract

Explicit chain-of-thought (CoT) reasoning substantially improves the reasoning ability of large language models (LLMs), but incurs high inference cost due to lengthy autoregressive traces. Existing latent reasoning methods offer a promis- ing alternative, yet they often treat reasoning as uniformly compressible, causing precision-critical intermediate steps to be overly compressed and thereby degrading reasoning accuracy. In this work, we propose Selective Latent Thinking (SLT), a reasoning paradigm that selectively compresses redundant reasoning spans into latent representations while preserving precision-critical spans as explicit CoT within the same reasoning trajectory. Our primary instantiation uses a lightweight decoder to anticipate a short upcoming reasoning span and applies confidence- based gating to determine the longest span that can be reliably compressed; the resulting selective compression policy is further optimized with trajectory-level reinforcement learning to balance answer correctness against reasoning cost. Ex- periments across four mathematical reasoning benchmarks show that which spans are compressed is critical: with the same compressor, compressing arbitrary spans reduces the GSM-Hard accuracy of Qwen3-4B from 54.7% to 17.3%, whereas keeping numerical tokens explicit retains 51.8%. SLT reduces explicit reasoning length by 58.4% with only a 2.8-point accuracy drop compared to explicit CoT on Llama-3.2-1B. Notably, the principle extends beyond this design: instantiated within existing latent reasoning models via a learned compression gate, it improves accuracy at equal reasoning length (+2.6 points over Latent-SFT) and at shorter length (+5.6 points over Latent-GRPO with 13% fewer tokens).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.