acceptodds
Under review as a conference paper at ICLR 2027

WASP: Weighted Asymmetric Sparse Attention to Accelerate both Prefilling and Decoding in Long-context Models

Abstract

Long-context inference stresses attention in two different ways: prefilling is compute-bound, while decoding is bound by memory bandwidth. Sparse attention methods mitigate this problem, but usually deal with only one of the two bottlenecks. Many of them also rely on attention patterns that exists in specific architectures, on probes that overfit the question-at-the-end format of common benchmarks, or on rigid top- budgets that fail to reflect how widely sparsity varies across layers, heads, tasks, and modalities. To this end, we introduce Weighted Asymmetric Sparse (WASP) attention, a training-free sparse attention designed to asymmetrically sparsify the two stages: block attention pools queries during prefilling and keys during decoding, Two primitives make WASP accurate: inverse density weighted (IDW) pooling, a TF–IDF-style summary that preserves the distinctive vectors that otherwise classical mean pooling would dilute, and an approximate Bernoulli sampler that adapts each head's budget to its own score distribution with no sorting. Across six text and multi-modal long-context benchmarks on Qwen3-8B, Llama-3.1-8B, and Qwen3-VL-8B, WASP matches the accuracy of dense attention on five of them while its attention layer runs faster at prefilling and at decoding, achieving an overall speedup that grows further with context length.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.