FReS: Factorized Recurrent State for Efficient Inference with Gated Delta-Rule Linear Attention
Abstract
Recurrent gated delta-rule linear attention has been widely adopted in recent hybrid language models, such as Gated DeltaNet (GDN) in the Qwen3 series and Kimi Delta Attention (KDA) in Kimi K3. By replacing the context-length-dependent key–value cache with a fixed-size recurrent state table, these architectures enable efficient long-context inference. However, these models maintains a dense recurrent state in each of its layers and heads for every active sequence. As the number of concurrent sequences increases, the replicated recurrent states lead to aggregate memory usage and traffic to grow linearly with concurrency, making GPU memory capacity and bandwidth critical bottlenecks for high-concurrency serving. To address these challenges, we propose FReS, a training-free inference method. Building on the observation that recurrent states exhibit low effective ranks, which indicates the potential for compression. FReS leverages low-rank matrix factorization to represent recurrent states compactly, reducing memory storage and traffic. FReS further employs an online, subspace-aware state maintenance to support incremental decoding while preserving most state insights and limiting error accumulation. Experiments show that FReS largely preserves dense-state quality while achieving up to the throughput of the dense-state baseline under high concurrency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.