Urchin: Parametrizing Attention for Long-Context Recall
Abstract
Transformers remain the dominant architecture for recall as they cache every key-value pair from context; however, the memory requirements for this cache limit scalability to ultra-long sequences. This has motivated work into finite-state relaxations of attention, such as linear attention and state-space models (SSMs). The test-time regression perspective offers insight into such layers by viewing attention as an online nonparametric regressor from keys to values, and proposing linear key-value regressors as parametric alternatives. Linear-state methods, such as Mamba-2, Gated DeltaNet, and Kimi Delta Attention, have seen widespread success; however, while competitive with attention on language modeling, they fall short on long-context recall. Recent dual-state SSMs such as Raven address this by introducing nonlinearity into the retrieval process, yet they sit outside the test-time regression framework, leaving it unclear how to extend or improve these methods in a principled manner. In this work, we view softmax attention as a nonparametric Gaussian mixture model over the key-value space, with a mixture component centered at each observed key-value pair. By parameterizing this density estimate with a fixed number of mixture components, this framework yields a design space of fixed-state models with the same softmax nonlinearity as attention. We unify past dual-state SSMs as particular instantiations of this framework, and use our framework to derive Urchin, a new sequence layer for long-context recall and language modeling. Urchin demonstrates strong recall across tasks and reaches near-perfect passkey retrieval on 1M-token contexts, a 500x extrapolation beyond its training length.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.