Understanding Evaluation Inconsistency in Diffusion Large Language Models
Abstract
Despite the capability of parallel decoding, diffusion large language models (dLLMs) require many denoising steps to maintain generation quality, motivating recent research on efficient decoding strategies. However, existing studies have reported inconsistent benchmark results even in seemingly identical evaluation settings, undermining the reliability of conclusions about dLLM decoding. To understand this issue, we perform a comprehensive empirical study of existing dLLM decoding methods across diverse evaluation protocols. Our analysis reveals that the performance gains of decoding methods are highly prompt-dependent, resulting in inconsistent rankings of methods across prompt templates. We further attribute this evaluation inconsistency to the high sensitivity of parallel decoding methods to prompt design, where even minor prompt variations can significantly alter generation behavior. Notably, we find that a well-chosen prompt template can yield markedly greater accuracy gains than simply increasing the decoding budget. Beyond prompt templates, our experiments indicate that other commonly overlooked evaluation settings can also affect the assessment of decoding methods. Based on these findings, we provide practical guidelines for reliably evaluating decoding methods for diffusion large language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.