acceptodds
Under review as a conference paper at ICLR 2027

Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs New Metrics

Abstract

Diffusion language models (DLMs) have emerged as the leading non-autoregressive alternatives to language modeling. Because their likelihoods are intractable, generative perplexity (gen-PPL) has become the de facto standard for reporting unconditional generation quality: the length-normalized likelihood of generated samples under a frozen autoregressive scorer, typically paired with a per-sample unigram entropy guardrail. We give an information-theoretic account of why this metric is unsound. Scoring model samples places the joint entropy of the generator inside gen-PPL, making collapse optimal for any fixed scorer. The reported per-sample unigram entropy ignores token order and cross-position dependence and therefore cannot correct this failure. Four training-free naive samplers reach or beat trained-model gen-PPL on LM1B and OpenWebText at non-degenerate unigram entropy, while producing text that is incoherent by construction. We then rebenchmark recent DLMs using distributional discrepancies and diversity diagnostics, and compare their rankings with LLM judgments. As part of this evaluation, we introduce a compression-based discrepancy that fits frozen byte-level coders to real and generated samples and measures local coverage and fidelity in both directions. The resulting metric profile provides a more informative view of unconditional text generation. We argue that DLM progress should be evaluated using a set of metrics that expose different properties of generated text.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.