acceptodds
Under review as a conference paper at ICLR 2027

HybriKV: KV Cache Eviction for Hybrid Linear-Attention Language Models

Abstract

Recent frontier language models increasingly adopt hybrid linear-attention architectures, replacing full attention with linear attention in most layers to reduce long-context memory overhead. However, even when retained in only a few layers, full attention still incurs KV caches that grow linearly with context length, posing a major memory bottleneck for long-context inference. KV-cache eviction provides a complementary way to alleviate this memory pressure, but existing methods are primarily designed for attention-only architectures and do not fully account for the distinct information pathways introduced by hybrid models. We identify two architecture-specific characteristics: the remaining full-attention layers exhibit less concentrated attention, and the linear pathway provides complementary token-level information beyond that available from the full-attention pathway. Motivated by these observations, we introduce HybriKV, a hybrid-aware KV-cache eviction method. To account for the less concentrated attention in the full-attention pathway, HybriKV aggregates evidence from multiple salient responses rather than relying primarily on the strongest one. To exploit complementary token-level information from the linear pathway, it reserves a small portion of the cache budget for positions highlighted by linear-layer signals. Across RULER, LongBench, and SCBench, HybriKV achieves near-lossless performance under approximately - KV-cache compression and reduces decoding-time attention latency by 45% compared with FlashAttention. It also consistently outperforms existing KV-eviction methods across hybrid models ranging from 7B to 48B parameters. Code is available at https://anonymous.4open.science/r/ICLR-Anonymous-397E.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.