ProbeRefine: Head-Adaptive Query Probing and Residual Refinement for Sparse Attention
Abstract
Large language models (LLMs) face substantial computational overhead during long-context prefilling due to the quadratic complexity of self-attention. Dynamic sparse attention alleviates this cost by selecting only a subset of tokens or blocks for computation, while proxy-based methods further reduce selection overhead by sharing block rankings across attention heads. However, head-specific budgets only determine how many blocks each head retains from a shared ranking, leaving limited flexibility in which blocks are selected when head preferences differ. In this paper, we introduce ProbeRefine, a training-free sparse attention method that preserves efficient shared proxy estimation while enabling head-specific block selection. ProbeRefine achieves this through two key designs: 1) Query-Guided Head-Specific Importance Estimation, which selects compact query probes for each head to estimate block importance and determine head-specific budgets; and 2) Gated Residual Block Replacement}, which revisits a bounded set of blocks beyond the shared cutoff and performs one-for-one replacement when a candidate receives sufficient head-specific importance and proxy support. Together, these designs allow different heads to refine their block selections without increasing the allocated block count or constructing full per-head attention maps. Experiments across Qwen2.5, Llama3.1, and Yi on RULER and InfiniteBench show that ProbeRefine consistently improves over ProxyAttn on both benchmarks and outperforms FlexPrefill across all three backbones on RULER while remaining competitive on InfiniteBench.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.