Enhancing Deep Residual LLMs with Depth-Recurrent Memory
Abstract
Residual connections have been a foundational design principle in large language models (LLMs), enabling the residual stream to accumulate layer-wise information across depth through unit-weight additions. Although this accumulation provides a stable information pathway, its fixed additive form offers limited explicit and input-dependent control over the evolving residual content. Recent studies have introduced constrained mixing or attention mechanisms to make residual information flow more adaptive. However, they still rely on the residual stream or stored intermediate layer representations as the carrier of cross-layer information, without an independent layer-wise memory and an explicit mechanism for controlling what to write, retain, and expose across depth. To fill this gap, we propose Depth-Recurrent Memory (DRM), a recurrent memory module whose state evolves along model depth. DRM maintains two fixed-width memory states per token, one for the attention pathway and one for the MLP pathway, and injects their gated readouts into the residual stream as additive updates. This provides explicit recurrent memory management while preserving the standard backbone sublayers and residual pathways. We instantiate the recurrent transitions with LSTM-style gating and co-design lightweight input projections with a parallel-residual implementation to control parameter and computational overhead and enable concurrent execution with backbone sublayers. Experimental results show that DRM outperforms the standard residual baseline on 16 of 21 benchmarks, with the strongest gains in math and code. In addition, the ablation studies validate the effectiveness of core components of DRM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.