acceptodds
Under review as a conference paper at ICLR 2027

Block-Level Positional Reconstruction for Sparse Long-context Prefilling

Abstract

Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length. Block-sparse attention offers a hardware-friendly alternative by selecting a small set of relevant key blocks for each query block, yet estimating block importance accurately without model retraining or costly token-level search remains challenging. Existing block estimators often pool RoPE-rotated token representations, causing within-block phase variation to induce frequency-dependent cancellation and distort position-sensitive block affinities. We empirically find that, although unrotated block summaries cannot exactly recover the omitted phase-aware interaction, its predictable component follows a frequency- and distance-conditioned response that transfers across input sequences. Based on this observation, we propose RePhase, which combines current unrotated block interactions with a reusable positional-response profile and the exact relative RoPE phase. The profile is fitted once from unlabeled forward-pass statistics without backpropagation or model-parameter updates, and enables direct block-level scoring at inference without token-level search. Experiments on long-context text and video tasks show that RePhase approaches full-attention performance across 4K–128K contexts, keeps block-estimation overhead below 3.4 ms, and achieves a 5.03 speedup over FlashAttention at 128K.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.