Deterministic Causal Multiscale Attention for Efficient Pretrained-Model Retrofitting
Abstract
Long-context transformer inference is dominated by quadratic causal attention, but many sparse methods trade efficiency for irregular execution or unclear model-level benefits. We present ButterFly, a deterministic, block-aligned causal multiscale attention pattern combining a small exact local neighborhood, the first global block, and logarithmically spaced historical blocks. For fixed block width, ButterFly uses O(n log n) selected token interactions and O(log n) graph communication depth per converted layer. Its static structure permits exact causal execution with predictable memory access and no learned router at inference time. We implement ButterFly in BF16 and retrofit eight of 28 attention layers in Qwen3-0.6B. On an NVIDIA RTX PRO 6000 Blackwell GPU, ButterFly crosses dense attention between 2K and 4K tokens, reduces selected interactions by 27.10× at 32K, and achieves 2.84×, 5.50×, and 11.91× per-layer prefill speedups at 8K, 16K, and 32K. The partial-model retrofit lowers warm whole-model time to first token by 7.7%, 11.3%, and 16.3% at the same lengths. Under a matched 1,000-step, 16.384M-token adaptation protocol, ButterFly ranks first across three training seeds, with mean report perplexity 1.68 versus 1.70 for an equal-budget static-random nonlocal graph, 1.69258 for dense continued training, and 1.81505 for equal-budget local sliding. These results establish ButterFly as a correctness-gated deterministic sparse-attention operator and selective pre-trained-model retrofit with measurable long-context systems value. They do not establish a fully subquadratic model, compressed total KV storage, or universal preservation of retrieval and reasoning capability; the remaining dense layers keep the complete model asymptotically quadratic.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.