LatentHalt: Reusing Termination Geometry for Adaptive Latent Reasoning
Abstract
Latent chain-of-thought (CoT) reduces reasoning cost by replacing textual rationales with continuous states, but fixed latent budgets leave the model without a per-problem stopping rule. We introduce , which reuses the termination geometry already present in an explicit-CoT model. In a supervised fine-tuned (SFT) 1B model, last-layer states that predict concentrate in a narrow angular cone, separable from other generation states with held-out AUROC . We freeze the cone's center and quantile radii, then train latent reasoning through a single joint curriculum. An auxiliary decoder supervises the content of each block, while a linear terminal hinge trains the final state of a fully latent sequence to enter the cone. At inference, the same region determines when to stop generating latent blocks and decode the answer, without a separately trained halting head. Across GSM8K-Aug, GSM-Hard, MultiArith, and SVAMP, all problems halt adaptively, using blocks on average. is on par with fixed-length supervised latent CoT on GSM-Hard and improves SVAMP and MultiArith accuracy by and points. It uses fewer estimated reasoning-phase FLOPs than explicit CoT, and its best model comes from a single -epoch run rather than the two-phase, -epoch CoconutSIM-CoT pipeline. The code is available at https://anonymous.4open.science/r/LatentHalt-EA19/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.