DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving
Abstract
Hybrid linear-attention architectures have recently scaled to large open-weight models, offering quality competitive with full attention while substantially reducing key/value (KV) cache growth. However, their in-place recurrent-state updates complicate cache management: prefix reuse requires state checkpoints alongside full-attention KV, while storing state checkpoints in full increases memory pressure, leading to more evictions and repeated prefill. By analyzing the decay structure of Gated DeltaNet (GDN) and Kimi Delta Attention (KDA), we find that different heads and channels retain prefix information over markedly different timescales, which we term retention horizons. This variation suggests substantial compression potential in persistent state checkpoints. Building on this observation, we introduce Decay-Aware State Compression (DASC), which derives retention horizons from model weights, packs selected long-horizon units into ragged state checkpoints, and optionally balances their storage across tensor-parallel ranks. On reuse, DASC either zero-fills omitted units or refreshes them from a bounded suffix with additional compute cost. We evaluate DASC on Kimi-Linear, Qwen3-Next, and Kimi-K3, covering model quality and serving performance. Across the evaluated retrieval and reasoning tasks, conservative configurations maintain quality close to full state checkpoint caching, while suffix refresh recovers accuracy at more aggressive compression. Implemented in SGLang, DASC provides – state checkpoint compression at representative operating points, reducing mean Time to First Token (TTFT) by 16.1–42.0% and mean end-to-end (E2E) latency by 4.9–10.2%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.