acceptodds
Under review as a conference paper at ICLR 2027

Triceps: Tri-Component Separable Attention for Long-Context Prefill

Abstract

The complexity of full attention during prefill in Transformer-based large language models is the main computational bottleneck in applications that require handling long contexts, such as long-document question answering, repository-level code agents or long book summarization. Existing sparse attention methods combine fixed patterns by union (attention sinks, sliding windows, retrieved blocks, vertical and slash patterns), resulting in a mask denser than the mass it captures requires, wasting computation on blocks that carry almost no attention. We instead show that the attention matrix admits a training-free multiplicative tri-component separable approximation: a relative-position factor, a key factor and a query-normalization factor. Combined multiplicatively, the three factors induce a score threshold that a union of fixed patterns cannot express. We dynamically evaluate this approximation per head to route computation between three modes: an long convolution when the error measured on the last query rows falls below a threshold, an sparse mode for low-density masks, and delegation to dense or another sparse attention technique otherwise. We evaluate this method on RULER and LongReason with two open-weight long-context models, Qwen3-30B-A3B-Instruct-2507 and Llama-3.1-Nemotron-8B-UltraLong-4M-Instruct, and, at extended context lengths (128k and beyond), we observe speedups in Time-To-First-Token (TTFT) over dense attention while achieving a superior speed–accuracy trade-off compared to existing sparse methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.