acceptodds
Under review as a conference paper at ICLR 2027

MaskAhead: Unifying KV Cache Selection and Eviction in Block Diffusion Language Models

Abstract

Block diffusion language models decode one block at a time against an exact key-value cache of the prefix, which grows with the prompt and with the model's own long generations. A small cache requires two decisions: which prefix entries the current block should read, and which entries later blocks may need. Selection is temporary. Eviction permanently removes entries from storage. Sparse-read methods cut repeated cache reads but keep the full prefix resident. Eviction reduces storage, but the decision is irreversible: entries removed now cannot serve later blocks. Block diffusion exposes masked positions in upcoming blocks before generation, so a no-write probe can obtain model-derived queries for retention without producing draft tokens. MaskAhead assigns the model's mask queries to the decisions they can inform: current-block queries rank temporary reads, while a no-write probe of upcoming blocks ranks the entries retained until the next eviction. Both decisions use an output-sensitive score that accounts for values. The surviving cache can also be stored at four bits. On MATH-500, Q-MaskAhead reduces peak KV by 75%, with a controlled accuracy degradation. For long-prompt QA, question-conditioned prefill reduces peak KV by 86.7% on average across HotpotQA, NarrativeQA, and MuSiQue.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.