Latent Reasoning Works: Recovering Chain-of-Thought Behaviour with Latent Thoughts
Abstract
Chain of thought (CoT) improves LLM reasoning, but requires a long, autoregressively generated trace, which is subsequently discarded: its value lies in its effect on the distribution over final answers; latent reasoning seeks that effect without the cost. Given the limited success of existing approaches, which do not achieve near-CoT performance at increased efficiency, we ask a more foundational question: do short latent contexts that provide near-CoT performance even exist? We demonstrate that they do: just a single latent token on HumanEval+ recovers 98% of CoT performance and 256 latent tokens even exceeds it on MATH-Hard; by contrast, the median CoT lengths on these datasets are over 1000 and 3000, respectively. We use knowledge distillation, distilling the effect of many CoT rollouts on answers into the latent context in a frozen LLM. Validation loss tracks training loss, indicating that the latents carry some concept of reasoning, not just memorisation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.