Learning to Echo Chain of Thought in Latent Space via Self-Distillation
Abstract
Chain-of-thought (CoT) reasoning has driven much of the recent progress of language models on complex problems, but generating reasoning traces autoregressively makes inference costly. Latent reasoning instead carries out intermediate computation in continuous hidden states, which can compress reasoning into far fewer steps. Supervising these states is the open problem: a text trace does not say what each state should encode, so existing methods rely on sparse answer supervision or on traces with clear reasoning-step boundaries, which most domains lack. We introduce Echo, which teaches a model to reason in latent space by *echoing* its own explicit reasoning. Our key insight is that as a model reads the CoT, its hidden states already trace the progression of its reasoning. Thus, randomly sampling and sorting them yields dense targets for latent states, without a separate teacher, a curriculum, and step annotations. Echo has two branches: an explicit branch that sees the question and CoT and a latent branch that sees only the question and learns to reason with a small set of latent slots. At each training step, the slots are supervised with freshly sampled, chronologically ordered hidden states from the explicit branch. With GPT-2 and Llama-3.2-1B-Instruct, Echo is the most competitive latent method in 9 of 12 settings across math, commonsense reasoning, and coding, even surpassing Explicit CoT in some cases. Since it refines a compact set of slots in parallel, it reduces inference latency by up to relative to Explicit CoT. These results suggest that models can effectively self-distill from the CoT hidden states and reason compactly in the latent space during inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.