acceptodds
Under review as a conference paper at ICLR 2027

Measuring Progress in Diffusion Language Model Pretraining

Abstract

In contrast to autoregressive models, it is an open question how to meaningfully compare training runs of diffusion language models (DLMs). In this paper, we propose an empirical protocol for choosing a performance metric for early DLM web pretraining, so that training runs can be compared by the time they take to reach a certain target. We ask three questions of each possible metric: 1) Does the metric rank training recipes robustly under slight changes to the performance target? 2) Does the metric consistently improve during DLM pretraining, so that several options for setting a target are possible? And 3) Does the metric recover the known rankings of autoregressive models, whose text likelihood can be evaluated directly? We apply this protocol to the following metrics from the DLM literature: accuracy in several downstream tasks, generative perplexity and bits per byte, distributional similarity, decoded-text validity, repetition and diversity. We do so across seven popular masked and uniform DLM training recipes. In our experiments, downstream performance provides the clearest stable target for tracking progress. Generative perplexity and bits per byte also give comparatively stable rankings and consistently improve, but every tested autoregressive reference model incorrectly ranks at least one pair from the GPT-2 family (even after entropy matching). MAUVE also improves persistently and gives stable local rankings, but reverses two GPT-2 comparisons. Unigram and bigram JS preserve the entropy-matched AR orders tested here, although their local DLM rankings change more often. Validity and repetition mainly change early in training or favor random text, so they do not provide natural stopping targets. Alongside diversity measures, they remain useful checks on the DLM's samplers. Based on our findings, we release an open-source speedrun built around downstream HellaSwag accuracy, one of the metrics that meets our checks most consistently here.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.