acceptodds
Under review as a conference paper at ICLR 2027

Blaze: Block-Level Adaptive KV Quantization Guided by Attention

Abstract

The KV cache is the memory bottleneck of long-context inference and grows with sequence length. Attention concentrates on a few positions, yet compression methods treat every token alike or evict tokens irreversibly. We present Blaze, a block-level adaptive KV cache quantization scheme that derives the precision of each block from the attention its positions receive. During prefill, Blaze marks a position closed when its attention stays below a threshold that decays with layer depth, since deeper layers attend to fewer positions, and it requires agreement across consecutive layers, so single-layer fluctuations do not drive the decision. Blocks are stored in INT8 or INT4 according to how many of their positions closed, blocks that carry long-distance positions keep the higher precision, and every decision is fixed during prefill, so decoding reads a lookup table and adds work per step. Across five models, Blaze compresses the KV cache by 70–71% at 4K and 8K prefixes at a perplexity within 1.6% of FP16, and it stays within 2 points of FP16 on eight downstream tasks and hard four-needle retrieval, where uniform INT4 falls 6 to 14 points behind. Blaze lowers the memory barrier for long-context inference by assigning precision to blocks instead of discarding tokens.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.