acceptodds
Under review as a conference paper at ICLR 2027

Ask the Masked Block: Predicting Denoised-State Relevance for KV Cache Eviction in Diffusion Language Models

Abstract

Diffusion large language models (dLLMs) incur substantial inference costs by repeatedly processing the full sequence with bidirectional attention. Temporal caching reduces redundant computation but introduces additional memory overhead by retaining full-sequence KV states. Spatial eviction reduces this cache footprint, yet current-state attention may not identify the entries needed after denoising. We introduce Ask-dLLM, a KV cache eviction framework that predicts denoised-state relevance conditioned on the current masked block. A lightweight scorer learns from relevance measured after the same frozen dLLM denoises the block and ranks external cache candidates once per block using masked-state representations. Across 23 reasoning, code, and long-context benchmarks, Ask-dLLM improves the overall mean score over the prior eviction baseline by 7.75 points on LLaDA and 2.39 points on Dream while retaining only 10% of external KV entries. Its overall mean score remains within 1 point of the unmodified Origin model on both backbones. On four LongBench datasets at an 8K context length, Ask-dLLM achieves approximately 9.7× and 11.8× the throughput of Origin on LLaDA and Dream, respectively, while matching or exceeding its average task score. Code is available at https://anonymous.4open.science/r/36675/README.md.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.