acceptodds
Under review as a conference paper at ICLR 2027

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models

Abstract

Recurrent LLM architectures have emerged as a promising approach for improving reasoning, as they enable multi-step computation in the embedding space without generating intermediate tokens. Models such as Ouro perform reasoning by iteratively updating internal representations while retaining a standard KV cache across iterations, causing memory consumption to grow linearly with reasoning depth. Consequently, increasing the number of reasoning iterations can lead to prohibitive memory usage, limiting the practical scalability of such architectures. In this work, we propose MELT, a novel architecture that decouples reasoning depth from memory consumption. Instead of using a standard KV cache per layer and loop, MELT maintains a single KV cache per layer shared across reasoning loops, while retaining per-loop KV caches for a small sliding window. This cache is updated over time via a learnable gating mechanism. To enable stable and efficient training under this architecture, we propose to train MELT using chunk-wise training to align training and inference. Empirically, we show that MELT models fine-tuned from pretrained Ouro parameters maintain the original model performance and outperform standard LLMs of comparable size, while maintaining a memory footprint comparable to those models and dramatically smaller than Ouro's. Overall, MELT achieves constant-memory iterative reasoning without sacrificing LoopLM performance, using only a lightweight post-training procedure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.