acceptodds
Under review as a conference paper at ICLR 2027

Shift-Accumulate Attention: Multiplier-Free Query-Key Products for Transformer Decoding

Abstract

Power-of-two quantisation turns a multiplication into a bit shift, but it has so far been applied primarily to the post-softmax attention–value product. We reformulate the earlier and larger score product, , by quantising the key cache to a signed power-of-two fixed-point representation, so that every scalar multiplication becomes a sign flip, a shift, and an integer accumulation. We call the resulting datapath ShiftAtten. A fixed-point headroom condition makes every shift a non-negative left shift, rendering the accumulation exact with respect to the quantised operands. A mantissa-extended power-of-two family then trades additional shift–add operations for accuracy across 4- to 8-bit representations, while a shift-exact online softmax extends the same treatment to the attention–value product. Fused CUDA kernels for a 1.1 B-parameter Llama decoder on an RTX 4090 reach the throughput of FP16 scaled dot-product attention at a 32k context length with a smaller KV cache. We separate the gain into from the narrower representation and from kernel engineering. On zero-shot reasoning, the 8-bit representation achieves percent versus percent for FP16, while at 4 bits it loses only percentage points compared with points for uniform INT4, despite worse perplexity and tensor-level error. A cost model calibrated by kernel disassembly shows that the remaining bottleneck is instruction issue rather than memory bandwidth: current GPUs provide a packed four-way integer multiply–accumulate but no corresponding packed shift primitive, and PTX scalar shift–accumulate is emulated. We therefore propose DS4A, a packed four-way shift-dot-accumulate instruction. Under the same kernel structure, DS4A is projected to reach FP16 throughput, compared with for an optimised INT8 MAC kernel, because the power-of-two representation moves 101 bytes per cached token versus 133 bytes for INT8.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.