Latent Recirculation: Refining Reasoning through Gated Latent Feedback
Abstract
Chain-of-thought (CoT) reasoning improves the performance of language models by making intermediate steps explicit, but decoding these steps autoregressively incurs substantial output overhead. Latent reasoning instead performs intermediate computation in continuous representations, offering a compact alternative to explicit thought generation. However, constructing useful latent updates and supervising their evolution remain challenging. We introduce Latent Recirculation, a lightweight framework that constructs a fixed, question-conditioned latent workspace and refines it through gated deep-to-shallow feedback. The feedback allows the reasoner to revisit the same workspace without increasing its length or generating intermediate text. To guide this refinement, an auxiliary explicit CoT branch provides a reference state during training, and a transition objective aligns latent updates with the displacement toward that reference. Both pretrained backbones remain frozen, with optimization restricted to the latent interface and low-rank adapters on the reasoner. Experiments on symbolic, arithmetic, and commonsense reasoning benchmarks show improved average accuracy across two reasoners, while ablations highlight the importance of combining feedback with transition supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.