ReMem: Reusable Latent Computation for Depth Extrapolation in Looped Language Models
Abstract
Looped Language Models (LoopLMs) provide a promising approach to test-time scaling by repeatedly applying parameter-shared Transformer blocks to increase latent computational depth without increasing model size. However, performance often deteriorates when recurrent depth extends beyond the regime encountered during training, limiting the effectiveness of recurrent-depth extrapolation. Several existing approaches seek to improve recurrent scaling through additional training or post-training, but require extra optimization. In this work, we revisit how information is propagated across recurrent iterations and find that previously computed recurrent states retain reusable computation that can support deeper recurrent processing. Based on this insight, we propose ReMem, a training-free recurrent memory mechanism that explicitly preserves and reuses historical computation during inference. ReMem maintains a compact, evolving recurrent history, allowing newly formed recurrent states to remain directly accessible to later iterations, and uses evidence from the frozen LM head to adaptively allocate historical contributions at the feature level. Importantly, this design requires no parameter updates or additional Transformer evaluations at matched recurrent depth. Experiments across language understanding, mathematical reasoning, and code generation benchmarks show that ReMem generally improves performance retention under recurrent-depth extrapolation, with gains becoming more pronounced as recurrent depth increases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.