ResMiSA: Residual Mixture of Sparse Attention for Long-Context LLM Inference
Abstract
Multi-head token scoring in the DeepSeek Sparse Attention (DSA) indexer remains costly for long-context inference. MISA reduces this cost by activating a subset of indexer heads as scoring experts, but omits the remaining heads' contributions and their direct gradient paths through the indexing score during training. Inspired by residual connections and shared experts, we introduce ResMiSA, which uses to decompose each head's response into an identity branch and a nonlinear residual branch. Because all indexer heads share the same key, their weighted identity branches merge exactly into one always-active shared expert that requires a single dot product per key. Routed residual experts supply nonlinear corrections, while the shared path retains a direct gradient path through every head. Applied without additional training to DeepSeek-V3.2 across 4K–128K contexts, one shared expert and seven residual experts replace 64-head scoring, with RULER scores only 0.06 percentage points below DSA on average and 0.21 points below at 128K. In frozen-backbone indexer warmup, ResMiSA achieves slightly lower attention-fitting loss and higher attention mass recall on training batches than DSA, using only or as many per-key scoring channels at the same trainable parameter count and training steps. Combined with LiteTopK, ResMiSA achieves a end-to-end time-to-first-token (TTFT) speedup over DSA at 1M context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.