Function-Valued Memory: Query-Dependent Values for Linear Attention
Abstract
Matrix-state linear attention compresses history into fixed-size recurrent memory. Its design involves both update dynamics and the Query-to-Value function class represented by the fast state; under a matrix-scale budget, expanding the latter becomes a capacity-allocation problem. We introduce Function-Valued Linear Attention (FVLA), which generalizes the fixed Value vector associated with each memory coordinate to a Query-dependent response function: the Query determines both how strongly the coordinate is read and what Value it returns. A separable quadratic parameterization learns response directions across contexts while adapting their Value coefficients online within each context. Although nonlinear in the Query, FVLA remains linear in its online coefficients, preserving shared-error Delta updates and chunkwise computation. Across Gated DeltaNet and Kimi Delta Attention experiments spanning 140M and 340M configurations, FVLA improves language modeling and in-context retrieval. Across the four host–scale pairs, one and four response modes yield median relative gains of 10.5% and 24.0%, respectively, in average accuracy over the six retrieval tasks. Ablations show that nonlinear responses use additional fast-state capacity more effectively than linear state expansion. Response-mode scaling yields a clear quality-state-compute trade-off, with overall modeling quality improving as more capacity is allocated to the response space. Models trained at 2K context retain language-modeling gains through 512K-token streaming evaluation and improve long-range retrieval beyond training length. Together, these results identify response-space allocation as a key design dimension for finite-state associative memory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.