acceptodds
Under review as a conference paper at ICLR 2027

ResMiSA: Residual Mixture of Sparse Attention for Long-Context LLM Inference

Abstract

Multi-head token scoring in the DeepSeek Sparse Attention (DSA) indexer remains costly for long-context inference. MISA reduces this cost by activating a subset of indexer heads as scoring experts, but omits the remaining heads' contributions and their direct gradient paths through the indexing score during training. Inspired by residual connections and shared experts, we introduce ResMiSA, which uses to decompose each head's response into an identity branch and a nonlinear residual branch. Because all indexer heads share the same key, their weighted identity branches merge exactly into one always-active shared expert that requires a single dot product per key. Routed residual experts supply nonlinear corrections, while the shared path retains a direct gradient path through every head. Applied without additional training to DeepSeek-V3.2 across 4K–128K contexts, one shared expert and seven residual experts replace 64-head scoring, with RULER scores only 0.06 percentage points below DSA on average and 0.21 points below at 128K. In frozen-backbone indexer warmup, ResMiSA achieves slightly lower attention-fitting loss and higher attention mass recall on training batches than DSA, using only or as many per-key scoring channels at the same trainable parameter count and training steps. Combined with LiteTopK, ResMiSA achieves a end-to-end time-to-first-token (TTFT) speedup over DSA at 1M context.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.