acceptodds
Under review as a conference paper at ICLR 2027

Sharing Memory Addressing Across Depth in Hybrid Gated Delta Networks

Abstract

Does every recurrent memory layer in a language model need its own learned addressing rule? We investigate this question by selectively sharing parameters across the depth of a hybrid Gated DeltaNet language model. Our primary intervention ties query and key projections and their short convolutions across recurrent layers, while preserving independent memory states, value and output transformations, feed-forward networks, and normalization parameters. A value-only sharing control removes exactly the same number of parameters, separating component choice from the amount of parameter reduction. In a 400M-class model trained on FineWeb-Edu, query/key sharing removes 4.40% of unique parameters and incurs a smaller held-out language-modeling penalty than parameter-matched value sharing at a common 10.22B-token checkpoint. Sharing additional update-control parameters has a similarly small penalty, whereas whole-block looping changes the quality–parameter trade-off substantially. However, QK-only tying reduces 4K single-needle retrieval from 80.2% to 37.6%, despite a 0.24% relative held-out perplexity increase; sharing update controls reaches 74.8%. These observations motivate a storage–semantic retrieval separation hypothesis: reusable memory addressing may coexist with depth-specific semantic selection and content transformations. We present this as a mechanistic interpretation to test, not an established functional decomposition. It requires neither identical address semantics across depth nor reduced expanded-layer computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.