AbsorbKV: Co-designing Selection and Writing for KV Cache Merging
Abstract
KV cache compression reduces the memory cost of long-context inference, but eviction irreversibly discards states that future queries may need. KV cache merging instead transfers information from evicted states into retained ones. However, selection constrains the effectiveness of merging by determining which retained states are available to receive information from evicted ones. Existing methods largely overlook this dependency, so a state may be evicted even when none of the retained states can faithfully preserve its information. We introduce *absorbability* to characterize, relative to a given set of retained states, whether the information of an evicted state can remain readable through a retained key without compromising the information originally stored in that retained state. Guided by this insight, we propose AbsorbKV, which co-designs a Hybrid Selector and Adaptive Absorption. The Hybrid Selector jointly considers importance and absorbability to retain states that are important or difficult to absorb. Adaptive Absorption then merges each evicted state only into a compatible retained state and learns how to adapt its information to the retained key through which future queries will retrieve it, while keeping retained keys unchanged. Extensive experiments demonstrate that AbsorbKV outperforms representative KV cache compression methods across diverse long-context tasks and models. On Qwen3-8B with a 10% KV cache budget, AbsorbKV achieves 42.10 on LongBench and 39.01 on RULER-16K, surpassing representative state-of-the-art KV merging methods such as AsymKV, which obtains 35.44 and 25.59, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.