Blaze: Block-Level Adaptive KV Quantization Guided by Attention
Abstract
The KV cache is the memory bottleneck of long-context inference and grows with sequence length. Attention concentrates on a few positions, yet compression methods treat every token alike or evict tokens irreversibly. We present Blaze, a block-level adaptive KV cache quantization scheme that derives the precision of each block from the attention its positions receive. During prefill, Blaze marks a position closed when its attention stays below a threshold that decays with layer depth, since deeper layers attend to fewer positions, and it requires agreement across consecutive layers, so single-layer fluctuations do not drive the decision. Blocks are stored in INT8 or INT4 according to how many of their positions closed, blocks that carry long-distance positions keep the higher precision, and every decision is fixed during prefill, so decoding reads a lookup table and adds work per step. Across five models, Blaze compresses the KV cache by 70–71% at 4K and 8K prefixes at a perplexity within 1.6% of FP16, and it stays within 2 points of FP16 on eight downstream tasks and hard four-needle retrieval, where uniform INT4 falls 6 to 14 points behind. Blaze lowers the memory barrier for long-context inference by assigning precision to blocks instead of discarding tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.