acceptodds
Under review as a conference paper at ICLR 2027

QuMix: Query-Guided Mixed-Precision KV-Cache Quantization for Block Diffusion Language Models

Abstract

Block diffusion lets diffusion language models (dLLMs) reuse completed blocks through a key–value (KV) cache, but storing and repeatedly reading this growing cache limits context length, batch size, and decoding speed. KV-cache quantization reduces these costs, yet has been studied almost entirely in autoregressive (AR) models. Applying an existing AR quantizer to a block-diffusion model, we find that two-bit quantization preserves nearly all of its original TriviaQA score, whereas reducing the precision to one bit causes a large drop. This motivates mixed precision, which lowers average storage by retaining more bits for sensitive components and using fewer bits elsewhere. Such allocation must preserve information that attention relies on, including the key of the attention sink—a token that many heads repeatedly attend to. AR models typically place this sink at the first token, making it possible to protect by position, but sink positions can vary in dLLMs. We therefore introduce QuMix, which uses queries to guide precision allocation without assuming a sink position. From calibration queries and keys, QuMix constructs a query-weighted basis and assigns more bits to directions whose errors have a larger effect on attention scores, without model training. To our knowledge, this is the first study of KV-cache quantization for block diffusion. QuMix retains 98–99% of the 16-bit LongBench score across three models at an average of 1.99 bits per cache element. A custom CUDA kernel that attends directly to the compressed cache supports a larger maximum batch and faster denoising steps at the same batch size (batch 15) than BF16 FlashAttention4 on an RTX PRO 6000 GPU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.