Depth with Memory: Token Dynamics of Attention Residuals
Abstract
Attention Residuals aggregate stored layer outputs with normalized depth weights. This operation generally produces a history-dependent continuous-depth equation, rather than a closed ordinary differential equation. Starting from a normalized Volterra formulation, we identify a common-reweighting condition under which the dynamics close and derive the exact coupled equations for token directions and radii, with query, key, and value matrices fixed across depth. For spherical causal attention with a symmetric positive-definite value matrix having a simple largest eigenvalue, we impose a block structure on the query–key product and an explicit condition ensuring positive radial forcing. These conditions yield positive radial bounds and global well-posedness. When each history rate has an infinite integral, we prove by causal induction that, outside a Lebesgue-null set of admissible raw initial embeddings, every token direction converges to one of the two leading eigenvector poles. The prescribed history kernels are held fixed as the initial embeddings vary. The radial equation asymptotically averages the leading eigenvalue multiplied by the signed attention mass. We give conditions for radial convergence and an explicit softmax formula for the limiting lengths, which may differ within one directional cluster. The conclusions concern a specified normalized model and a sufficient clustering regime, not an arbitrarily trained Attention Residual networks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.