HybriKV: KV Cache Eviction for Hybrid Linear-Attention Language Models
Abstract
Recent frontier language models increasingly adopt hybrid linear-attention architectures, replacing full attention with linear attention in most layers to reduce long-context memory overhead. However, even when retained in only a few layers, full attention still incurs KV caches that grow linearly with context length, posing a major memory bottleneck for long-context inference. KV-cache eviction provides a complementary way to alleviate this memory pressure, but existing methods are primarily designed for attention-only architectures and do not fully account for the distinct information pathways introduced by hybrid models. We identify two architecture-specific characteristics: the remaining full-attention layers exhibit less concentrated attention, and the linear pathway provides complementary token-level information beyond that available from the full-attention pathway. Motivated by these observations, we introduce HybriKV, a hybrid-aware KV-cache eviction method. To account for the less concentrated attention in the full-attention pathway, HybriKV aggregates evidence from multiple salient responses rather than relying primarily on the strongest one. To exploit complementary token-level information from the linear pathway, it reserves a small portion of the cache budget for positions highlighted by linear-layer signals. Across RULER, LongBench, and SCBench, HybriKV achieves near-lossless performance under approximately - KV-cache compression and reduces decoding-time attention latency by 45% compared with FlashAttention. It also consistently outperforms existing KV-eviction methods across hybrid models ranging from 7B to 48B parameters. Code is available at https://anonymous.4open.science/r/ICLR-Anonymous-397E.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.