acceptodds
Under review as a conference paper at ICLR 2027

HiServe: Low-Rank Prefix Caching for Hybrid LLMs

Abstract

Hybrid language models combine recurrent layers with full attention to achieve competitive model quality while reducing the growth of the key/value (KV) cache. However, reusing shared prefixes requires separate recurrent-state checkpoints at cached boundaries. Storing these checkpoints in full increases memory pressure, eviction, and repeated prefill. Across five recurrent and hybrid language models, we find that recurrent states exhibit approximately low-rank structure, enabling substantial compression through truncated SVD. HiServe performs this compression asynchronously on the GPU and reconstructs states on reuse. At rank 16, HiServe provides a 3.98× reduction in same-precision storage for the evaluated Qwen3.5 recurrent-state matrices, or 3.73× including unchanged convolution states. Separate evaluations, including long-context tasks and periodic recompression, show benchmark-score changes from -0.63 to +2.00 percentage points. On multi-turn ShareGPT reference-history replay with Qwen3.5-4B on an NVIDIA GB10, HiServe improves throughput by 11.5% and reduces mean time to first token by 47.9% over SGLang's uncompressed prefix cache at a matched 15.23 GB cache budget. Gains depend on cache allocation and budget. Complementary trace-driven simulations with Marconi show a mean token cache-hit-rate improvement of 20.0 percentage points on ShareGPT under equal cache-memory budgets. Low-rank recurrent-state compression can thus increase prefix reuse and improve serving performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.