Sparse Attention with Entropy-Triggered KV Repair
Abstract
Sparse attention reduces the per-step decoding cost of language models trained with full attention. However, it can lengthen generations and lower accuracy, offsetting the per-step savings with additional tokens. We find that sparse-attention errors are most detrimental at steps where the full-attention next-token entropy is high: applying sparse attention only at these steps degrades accuracy far more than applying it at an equal number of random steps. We further find that at these high-entropy steps, most of the degradation stems from errors accumulated into the key-value (KV) cache during earlier sparse decoding, so even attending to the entire cache at such steps results in performance degradation. Motivated by these findings, we propose **HERA** (**H**igh-**E**ntropy-triggered KV **R**ep**A**ir), a method which uses sparse next-token entropy to trigger a KV repair — a full-attention pass that recomputes and replaces the cached KV states to yield the dense next-token distribution for sampling. We evaluate **HERA** on three reasoning models across four benchmarks. On MATH-500 with Qwen3-8B, **HERA** improves accuracy by 18.8 percentage points over the sparse-decoding baseline (i.e., Quest) with 25.7% fewer tokens, and by 5.6 points over a KV repair at fixed intervals baseline (i.e., ReSA) with 7.4% fewer tokens and 13.7% fewer repair calls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.