POEM: Output-Relative Block-Sparse Attention for Long-Context Prefilling
Abstract
As large language models (LLMs) evolve to support increasingly long contexts, accelerating prefilling without compromising output accuracy becomes a critical challenge. Block-sparse attention addresses this by retaining only the most important key-value blocks, but existing importance proxies often discard blocks and cause output errors. These proxies generally estimate a block's contribution to the dense attention output (i.e., attention-weighted value), implicitly assuming that removing the block loses exactly this contribution. In this work, we show that this assumption does not hold because attention renormalization amplifies the retained blocks' contributions. As a result, the output error depends not only on the removed attention mass but also on the distance between the removed block's output and the dense output. We call the distance output deviation and empirically find that it varies by over 20 even among blocks with comparable attention mass, so existing proxies often discard blocks whose large output deviation leads to large output errors. To address this gap, we propose Prefilling via Output Error Minimization (POEM), a training-free method that combines each block's contribution and output deviation into an output-relative score, reducing the output error of each head. To connect the head-level optimization to the model's predictions, we derive a bound on the prediction divergence that weights each head's output error by its sensitivity, which guides POEM to calibrate the retention threshold of each head. On long-context benchmarks, POEM maintains accuracy parity with dense attention while achieving up to a attention speedup over FlashAttention-2.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.