MassAlloc Attention: Let Attention Allocate Its Own Compute
Abstract
Long-context full softmax attention (FullAttn) often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MassAlloc Attention (MALA), a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocate post-score computation. Forward uses its evolving online-softmax normalizer, while backward reuses the finalized normalizer to derive nested retained support using only standard attention state. A common tolerance governs training and inference, allowing the retained work to adapt across queries, heads, layers, and inputs. MALA reduces low-contribution post-score computation while retaining quadratic QK score discovery. A matched-work study at 8K isolates the benefit of distribution-adaptive allocation: under exactly matched total post-score work, MALA approaches a per-instance reference-mass oracle, with mean omitted mass of 0.0188% versus 0.0182%, while static allocations perform substantially worse. Across context lengths from 1K to 32K tokens, the same tolerance maintains low output and gradient errors relative to the reference operator. Across a broader controlled associative-recall comparison, MALA closely tracks FullAttn as context grows, reaching 89.67% accuracy at 8K compared with 89.97% for FullAttn. In an attention-operator benchmark at 128K tokens on 8 H100 GPUs with tensor parallelism, MALA reduces forward and backward latency during training by and and decoding latency during inference by relative to FullAttn, while retaining FullAttn-level per-rank peak operator memory. Across scaling-law training from 0.6B to 14B parameters on 128 H100 GPUs, MALA closely tracks FullAttn in perplexity while reducing total training FLOPs, with a 23.1% reduction at 14B during 32K-context training. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results indicate that allocating post-score computation according to normalized attention contributions can retain the evaluated capabilities of FullAttn while reducing attention computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.