RBS-ATTENTION: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
Abstract
Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection reduces this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose RBS-Attention, a training-free sparse-prefill method with two complementary branches. A centroid base branch captures average relevance, while a rescue branch uses the maximum key-block radius and its prompt-, layer-, and head-dependent distribution to identify blocks at risk of underestimation. Independent thresholds and a mask union focus rescue on dispersed blocks while preserving block-sparse FlashAttention execution. On H100 GPUs, RBS-Attention achieves standalone prefill-attention speedup and end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8. On Qwen3-32B, it obtains 88.65 overall RULER accuracy versus 89.52 for dense attention; LongBench-v2, InfiniteBench, and Video-MME provide additional quality evaluation. Matched-density and selector studies further support radius-adaptive rescue as an effective approach to long-context prefill.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.