MaskAhead: Unifying KV Cache Selection and Eviction in Block Diffusion Language Models
Abstract
Block diffusion language models decode one block at a time against an exact key-value cache of the prefix, which grows with the prompt and with the model's own long generations. A small cache requires two decisions: which prefix entries the current block should read, and which entries later blocks may need. Selection is temporary. Eviction permanently removes entries from storage. Sparse-read methods cut repeated cache reads but keep the full prefix resident. Eviction reduces storage, but the decision is irreversible: entries removed now cannot serve later blocks. Block diffusion exposes masked positions in upcoming blocks before generation, so a no-write probe can obtain model-derived queries for retention without producing draft tokens. MaskAhead assigns the model's mask queries to the decisions they can inform: current-block queries rank temporary reads, while a no-write probe of upcoming blocks ranks the entries retained until the next eviction. Both decisions use an output-sensitive score that accounts for values. The surviving cache can also be stored at four bits. On MATH-500, Q-MaskAhead reduces peak KV by 75%, with a controlled accuracy degradation. For long-prompt QA, question-conditioned prefill reduces peak KV by 86.7% on average across HotpotQA, NarrativeQA, and MuSiQue.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.