acceptodds
Under review as a conference paper at ICLR 2027

Understanding Evaluation Inconsistency in Diffusion Large Language Models

Abstract

Despite the capability of parallel decoding, diffusion large language models (dLLMs) require many denoising steps to maintain generation quality, motivating recent research on efficient decoding strategies. However, existing studies have reported inconsistent benchmark results even in seemingly identical evaluation settings, undermining the reliability of conclusions about dLLM decoding. To understand this issue, we perform a comprehensive empirical study of existing dLLM decoding methods across diverse evaluation protocols. Our analysis reveals that the performance gains of decoding methods are highly prompt-dependent, resulting in inconsistent rankings of methods across prompt templates. We further attribute this evaluation inconsistency to the high sensitivity of parallel decoding methods to prompt design, where even minor prompt variations can significantly alter generation behavior. Notably, we find that a well-chosen prompt template can yield markedly greater accuracy gains than simply increasing the decoding budget. Beyond prompt templates, our experiments indicate that other commonly overlooked evaluation settings can also affect the assessment of decoding methods. Based on these findings, we provide practical guidelines for reliably evaluating decoding methods for diffusion large language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.