Not All Tokens Deserve Equal Resolution: Score-Adaptive Indexing for Sparse Attention in Long-Context LLMs
Abstract
Indexer-based sparse attention reduces the cost of the main attention operator by retrieving only a small subset of the context, but the indexing stage itself is expensive for long sequence lengths. Hierarchical indexing reduces this cost by representing contiguous regions with compact summaries and refining only selected candidates. However, fixed-size partitions assume that every region can be represented equally well by one summary. While mean pooling is reliable when nearby keys are similar, it can obscure important token-level variation when they are not. We introduce SCALE (Score-Adaptive Indexing for Long-Context LLMs), a training-free hierarchical index that improves summary fidelity under a fixed budget by adapting which tokens are grouped together. For a fixed chunk, the mean key minimizes squared reconstruction error. For the DSA scoring function used in our evaluation, we further show that key dispersion controls an upper bound on the score approximation error induced by mean pooling. SCALE therefore splits high-dispersion regions into more chunks and merges homogeneous neighbors, while retaining the original mean representation and token-level indexer for exact refinement. Using the same number of chunks, SCALE improves Top- candidate recall over fixed-block indexing by – percentage points while consistently reducing score approximation error. With one summary per tokens on average, SCALE-64 improves pooled RULER accuracy from to and LongBench-v2 accuracy from to over fixed-block HISA-64, a recent strong fixed-block hierarchical baseline. On the 128K serving workload, the same configuration reduces end-to-end latency by relative to flat DSA (DeepSeek Sparse Attention), from s to s, while matching DSA on LongBench-v2 () and remaining close on RULER ( versus ).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.