acceptodds
Under review as a conference paper at ICLR 2027

FactorState: Dynamic Low-Rank Tracking for Memory-Efficient Gated Delta Networks

Abstract

Long-context inference incurs growing memory costs from attention key–value (KV) caches. Gated Delta Networks (GDNs) mitigate this growth by summarizing history in fixed-size recurrent matrix states rather than maintaining a growing KV cache. Yet in hybrid large language models, these states are replicated across layers, heads, and concurrent requests, leaving a substantial aggregate memory footprint. Existing work has revealed low-rank structure in recurrent states and used it to motivate compression, but rank alone does not determine whether retained subspaces remain useful during decoding. We identify a temporal phenomenon, capacity–geometry decoupling: at a fixed rank, optimal approximations of the current state maintain nearly unchanged normalized squared reconstruction error, while reusing earlier subspaces leads to growing reconstruction and memory-read errors. Motivated by this observation, we introduce FACTORSTATE, a training-free representation that tracks the evolving recurrent state with compact factors under a fixed memory budget. FACTORSTATE performs recurrent reads and updates directly in factor space and periodically recompresses the factors, leaving pretrained weights and dense prompt processing unchanged. Experiments across Qwen and OLMo models show that FACTORSTATE reduces recurrent-state storage by approximately 61% while achieving task performance close to dense execution. GPU evaluations further demonstrate lower execution time and peak memory under both fixed-active-batch and continuous-batching workloads.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.