Cache Complete Cone Prefill for Local-Attention Decoders
Abstract
Prefill processes a prompt before generation begins and contributes directly to the time to first token. We propose *Cache Complete Cone Prefill* (C3P), a training- and calibration-free accelerator for local-attention decoders that retains pretrained weights and the native decoding interface. C3P traces dependencies backward from the requested logits and the cached state read by later decoding, then executes only the resulting *cone*. We prove that it reproduces both exactly in real arithmetic and is minimal among schedules that only delete original operations. We derive closed-form expressions for how many token positions each layer must process, explaining how prompt length, attention window size, and model depth determine the available savings. On an A100, C3P reduces BF16 Mistral-7B prefill latency by up to 10.6% over a FlashAttention-2 baseline that already uses final-layer slicing. On synthetic FP32 local decoders, it achieves up to speedup and 80.30% lower peak allocated memory over final-layer slicing. We further ask what information must be retained to preserve the prompt logits and future outputs when the decoding state is encoded differently. For a single attention layer with fixed keys, we determine the minimum information that must be retained from cached values to reproduce future outputs exactly, and how this requirement depends on the number of decoding steps. For purely local decoder stacks, we establish conditions under which the prompt logits and the logits of a sufficiently long continuation together uniquely determine the input vectors at every prompt position in the cone, so re-encoding cannot discard the information in those inputs. In sum, C3P combines training-free, plug-and-play prefill acceleration with rigorous guarantees on exactness and computational minimality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.