Depth Memory
Abstract
Residual connections propagate information over depth by additively accumulating layer outputs. Recently proposed alternatives, such as Manifold-Constrained HyperConnections (mHC) and Attention Residuals (AttnRes), instead use more expressive depth aggregation mechanisms that improve language modeling performance at scale. In this work, we consider the mixing of representations over depth a problem of storing and retrieving information across layers. Building on this perspective, we propose Depth Memory (DM), a framework that interprets the standard transformer, as well as mHC and AttnRes, as propagating information forward using different streams. One such stream can be a memory stream, which is queried to enrich the input to every layer, and whose output is subsequently written back into the memory for use by later layers. For AttnRes, this memory is a growing cache of explicit layer representations. For mHC, it is an implicit representation stored in an expanded stream. We then instantiate DM as a Linear Depth Memory (LDM), which propagates forward three streams: a standard residual stream, a memory stream, and an inverse covariance stream. We formulate the dynamics of the memory stream as a depth regression problem that is solved online via recursive least squares. We furthermore propose a hardware-aware implementation of LDM that reduces peak activation memory and memory traffic, enabling high-throughput training and inference. LDM consistently achieves the lowest perplexities among the evaluated PreLN, mHC and AttnRes baselines at the 400M and 1.5B parameter scales, alongside competitive zero-shot accuracy, showcasing the practical utility of our framework.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.