Finding the Needles in Long-Context Attention through Sparse Features
Abstract
In long-context decoding, high attention scores often concentrate on a small fraction of historical tokens scattered throughout the context. Attending only to these tokens can preserve generation quality, but finding them efficiently is a needle-in-a-haystack problem. We show that these tokens are not arbitrarily distributed in feature space: although they may differ in many respects, they share query-relevant features in a high-dimensional sparse representation. To expose this structure, we develop a shared query-key sparse autoencoder (SAE) that decomposes dense queries and keys into a common dictionary of sparse features. Most high-scoring keys share active features with the query, whereas such overlap is uncommon among the remaining keys. Building on this finding, we introduce PrismKV, which indexes historical keys by their active SAE features to form overlapping buckets. Compact bucket summaries identify relevant tokens without scanning the entire history, allowing attention to operate only on the selected original KV entries. Across RULER and LongBench on the 27B model, PrismKV retains an average of 99.90% of full-attention performance. Relative to optimized full attention, it achieves attention-path speedups of 3.37x on Qwen3-8B at a 128K context length and 2.78x on the 27B model with a 256K total window.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.