MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
Abstract
Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length. This flexibility, however, incurs quadratic computational complexity during training and a key-value cache that grows linearly during autoregressive inference. Recurrent alternatives offer efficient decoding by compressing the entire history into a fixed-size state, but often underperform on recall-intensive tasks because earlier associations are usually overwritten by subsequent updates, leaving only the most recent contextual information available. In this paper, we introduce Memory-Anchor Routing across Context History (MARCH), a network architecture that effectively scales state-space models beyond a fixed-size dimension while maintaining computational efficiency over long sequences. MARCH periodically caches cumulative recurrent-state checkpoints as state anchors and associates each anchor with a compact, content-conditioned anchor key. This allows MARCH to maintain a memory bank that grows with context length, providing a controllable trade-off between historical resolution and memory cost. At each token, MARCH produces an anchor query to attend to all causally available state anchors and computes the output by aggregating the historical-anchor readouts together with the current-state readout in an attention-style manner. We show that, after standard pretraining, MARCH consistently outperforms multiple linear-attention variants across commonsense reasoning, LongBench, and in-context retrieval. These results demonstrate that content-routed state caching substantially strengthens recurrent long-range memory while preserving its native computational path.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.