acceptodds
Under review as a conference paper at ICLR 2027

Step-Wise Endorsement: Evaluating Approximate KV Caching in Diffusion LLMs One Decision at a Time

Abstract

Most diffusion large language models (dLLMs) use bidirectional attention, so an exact KV cache is impossible: every newly decoded token changes the keys and values (KV) of all other tokens. Approximate KV caching methods such as Fast-dLLM reuse KV from an earlier step and recompute (refresh) them only occasionally. Such a method is good to the extent that the cached model approximates cache-free exact decoding, a property no metric measures directly. The refresh interval, however, orders these caches in advance: refreshing more often leaves less of the decoding on stale KV. Varying only this interval gives a refresh-dose scale of caches whose order is known, which any sound evaluation must follow. Accuracy, the evaluation used most, does not: in two of four configurations of Fast-dLLM, and on HumanEval and MBPP, it calls the cache that never refreshes the better one. Comparing the cached and the exact output token by token barely separates the settings. We propose step-wise endorsement: before each step of the cached model, the exact model receives the same state and decides there itself, and four metrics record how often and how far the two disagree on which [MASK] to decode and which token to write. All four follow the scale in every configuration on math, at both ends of it on code, and along the refresh knobs of dLLM-Cache, dKV-Cache, and Elastic-Cache; with a tenth of the questions, they order the settings far more reliably than accuracy does with all of them. They also show why Fast-dLLM works: stale KV change decisions, but rarely, and mostly where the exact model itself is uncertain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.