LaTER: Adaptive Latent Computation for Efficient Explicit Reasoning
Abstract
Chain-of-thought (CoT) reasoning realizes additional inference-time computation through autoregressive token generation, while continuous latent reasoning removes this dependence on discrete intermediate tokens. However, making the entire derivation latent can degrade reasoning accuracy. We propose **La**tent-**T**hen-**E**xplicit **R**easoning (**LaTER**), which uses a continuous latent prefix to prepare a shorter explicit derivation. LaTER recurrently feeds projected hidden states back as inputs, then transitions to token generation while retaining the latent key-value cache. We study both a training-free variant, which determines this transition from vocabulary-probe entropy and terminating-token predictions, and a supervised variant trained on Latent-Switch-69K, a corpus of 69,745 examples that we construct by pairing solution intuitions with shortened derivations. To learn reliable stopping behavior, we introduce a halting objective that accounts for a training-inference mismatch in which forced training rollouts can mask premature stop predictions. On Qwen3-14B, supervised LaTER improves accuracy while reducing reasoning length on all seven evaluated benchmarks relative to both standard CoT and a same-data CoT-SFT control. On AIME 2025, accuracy increases from 70.0% to 80.0%, while latent steps plus explicit tokens decrease from 15,730 to 10,575. On Qwen3-8B, wall-clock decoding time decreases by 19.1% on average across benchmarks in our implementation. These results suggest that latent and explicit reasoning are complementary, and that the transition between them is a key design choice for improving the accuracy-efficiency tradeoff. Our code, data, and model are available at https://anonymous.4open.science/r/LaTER-EF4D.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.