Lost in a Few Heads: Diagnosing and Repairing Block Selection in Sparse Attention Language Models
Abstract
Sparse attention reduces the compute and memory cost of long-context inference, but already-trained sparse attention models may lose substantial accuracy at their deployed selection budgets. We localize where this loss occurs. Using exact at- tention mass selection as an oracle diagnostic per token and per head, we locate the multi-key retrieval loss in the InfLLM-V2 family to three or four of its 64 key-value heads, the same heads in four checkpoints of one backbone. These blind heads attend to the needle, yet their selector misses it likely because the mean-pooled key summary scored by the selector dilutes single decisive token. A retrieval-head score ranks heads whose selector already finds the needle above the blind heads. Once the heads are found, the loss can be repaired without training most simply by using full attention in them, costing the least for a single sequence. An alternative is a scorer ranking blocks with each query head’s full softmax over a 4- bit copy of the keys while keeping the shipped budget and reading an eighth of the bytes that full attention reads in blind heads. On prompts held out from every design we test, the two repairs recover 31.4 and 30.1, respectively, out of the 33.2 points on RULER MK3 multi-key retrieval that InfLLM-V2-8B loses, and 41.0 and 41.1 of 41.8 on MiniCPM4.1-8B at 64K. On four other selector designs, the scorer recovers most of the loss by acting in many or all of their heads.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.