From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers
Abstract
Chain-of-thought (CoT) reasoning often improves language-model performance by giving models additional computation before answering. However, explicit CoT expresses this computation as a sequence of autoregressively generated tokens. Latent reasoning replaces these tokens with compact continuous states, but most autoregressive latent-reasoning methods retain a left-to-right dependency among latent vectors. We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace. A shared transformer is reapplied for a small number of refinement iterations, jointly updating the latent slots based on the prompt and the evolving workspace state. Using the refined state, a probabilistic head predicts a distribution from which latent tokens are sampled in parallel and used to condition an autoregressive decoder for answer generation. Training uses continuous representations derived from explicit CoT together with a final-answer prediction loss and likelihood-based supervision of the latent states. Across HumanEval, HumanEval+, MBPP, and MBPP+, the method achieves the highest four-benchmark mean among the evaluated methods, performing on par with Reasoning SFT while outperforming the base model, answer-only SFT, and NF-CoT. Furthermore, compact latent states and parallel slot sampling reduce the serial thought-generation bottleneck of explicit CoT and other autoregressive methods while preserving probabilistic latent modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.