HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing
Abstract
We introduce **HySparse2**, a hybrid sparse attention architecture with **two-level KV sharing**. At the outer sharing level, **KV Bridging** adopts a YOCO-style structure with a self-decoder and a cross-decoder, but bridges only full-attention layers. The self-decoder uses hybrid sliding-window attention, while the cross-decoder uses hybrid sparse attention. The KV caches for full-attention layers in the cross-decoder are generated from the hidden states of full-attention layers in the self-decoder. At the inner sharing level, HySparse2 retains HySparse's core **KV Reuse** design with two refinements. First, it replaces block-level sparsity with *token-level* sparsity to enhance long-horizon agentic capabilities. Second, it removes the separate SWA branch from sparse layers and instead forces a sliding window of recent tokens into the sparse selection. This two-level KV sharing mechanism enables HySparse2 to construct all KV caches from self-decoder hidden states. During inference, prefill can therefore exit early after the self-decoder, skipping all cross-decoder layers. After posttraining, HySparse2 outperforms HySparse and Hybrid SWA on long-context retrieval and multi-turn agentic tasks, while substantially reducing inference cost and KV-cache size.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.