CSAttention: Centroid-Scoring Sparse Attention for Reusable-Prefix Long-Context LLM Serving
Abstract
Long-context LLM serving increasingly relies on long, reusable prefixes, such as cached retrieval contexts, tool schemas, agent memory, and domain-specific system prompts. In this write-once, read-many regime, dense attention over the accumulated KV cache becomes a dominant decode-time bottleneck, while aggressive sparse attention often loses accuracy because online queries and cached keys follow different distributions. We propose CSAttention, a training-free sparse attention method designed for reusable-prefix long-context inference. CSAttention shifts expensive retrieval work from online decoding to an offline prefix-build stage: it clusters prefill queries in subspaces, precomputes centroid-to-key partial scores, and stores bounded Top- lookup tables for each query centroid. At decode time, each new query selects its nearest query centroids, fetches the corresponding shortlists, aggregates centroid scores by key index, and performs attention only over the selected KV entries. This formulation does not remove the linear memory growth of long-context KV cache; instead, it creates a decode-stationary auxiliary index whose capacity is fixed after a reusable prefix is built and whose online work is regular and bounded by the shortlist budget. Extensive experiments on LongBench and LongBench-v2 show that CSAttention preserves near-full-attention accuracy at 95% sparsity across Llama, Qwen, and Mistral backbones. Compared with strong sparse-attention baselines, CSAttention achieves higher recall and downstream accuracy under matched sparse budgets, and delivers up to decode-throughput speedup and up to end-to-end speedup in long-prefix, moderate-to-long generation regimes, while making its TTFT, memory, and batch-scaling trade-offs explicit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.