acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Adaptive Sparse Attention over a Multi-Resolution Cache Index

Abstract

Long-context inference in large language models is increasingly constrained by the cost of reading a growing key-value (KV) cache. To address this bottleneck, we introduce Hierarchical Adaptive Sparse Attention (HASA), a sparse attention operator that adapts how much detail it retrieves from different parts of the context. HASA organizes each layer’s cache into a hierarchy of token-level representations and learned summaries over progressively larger segments. For each query and layer, a learned router selects the resolution at which each historical span enters attention, ranging from token-level KV states to increasingly compact learned summaries. The summaries and router are learned through distillation and task-level fine-tuning with the backbone frozen. On Qwen3-8B, HASA improves the average score across all sixteen English LongBench tasks by +1.1 points over full attention while attending to 4.3× fewer KV entries. With a fused decode kernel, HASA decodes faster than FlashAttention-3 at all measured context lengths, reducing latency by 12.3% and 20.0% at 65K and 131K, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.