acceptodds
Under review as a conference paper at ICLR 2027

LoLa: Looped Latent Reasoning via Hidden State Feedback

Abstract

Large reasoning models (LRMs) have achieved strong performance on reasoning tasks through chain-of-thought (CoT) generation, yet the resulting trajectories are often lengthy and redundant, leading to substantial inference overhead. Recent work has explored latent reasoning to improve efficiency, but existing methods typically either generate latent representations in a single forward pass without iterative refinement or build recurrence into the model during large scale pretraining. We propose looped latent reasoning (LoLa), a framework that combines a lightweight latent reasoner with a frozen base large language model (LLM). The reasoner maintains a latent state and iteratively updates it using hidden state feedback from the base LLM. The final state is projected into latent thought tokens that condition answer generation. We train the reasoner using supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO). We evaluate LoLa with two base LLMs on five reasoning benchmarks. With DeepSeek-R1-Distill-Qwen-1.5B as the base model, LoLa outperforms existing efficient reasoning methods. With Qwen3-4B, LoLa improves pass@1 over its non-thinking mode from 54.07% to 58.64% and pass@4 from 65.78% to 72.36%, while also reducing inference latency. Once trained for a given base LLM, LoLa provides a plug-and-play reasoning mode that can be enabled on demand without modifying the base model, while explicit CoT remains available for harder problems. Code is available at https://anonymous.4open.science/r/loopedlatent_iclr2027.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.