acceptodds
Under review as a conference paper at ICLR 2027

MISA:Mixture of Indexer Sparse Attention for Long-Context LLM Inference

Abstract

DeepSeek Sparse Attention (DSA) uses a learned multi-head indexer to score prefix tokens and select the top- tokens for sparse attention, but indexing becomes costly at long context lengths. We propose **MISA** (**M**ixture of **I**ndexer **S**parse **A**ttention), a training-free, drop-in replacement that treats indexer heads as a pool of experts. A lightweight router uses block-level statistics to select a query-dependent subset of active heads for token-level scoring, reducing the per-query indexing cost from to , where . A hierarchical variant, MISA, further re-ranks an enlarged candidate set with the original DSA indexer. On the 13-subtask RULER benchmark across six context lengths from 4K to 128K, MISA achieves an average score of 94.94, second only to DSA at 95.46. End-to-end prefill evaluation shows that MISA reduces mean time to first token by a factor of on the full model at 256K and on a pruned 8-layer model at 512K. These results demonstrate that head-axis routing provides a practical quality–efficiency trade-off for fine-grained sparse attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.