acceptodds
Under review as a conference paper at ICLR 2027

Sampling More, Getting Less: A Token-Level Diagnostic for LLM Generation Diversity

Abstract

Diversity is essential for language-model applications ranging from creative generation to scientific discovery. Yet modern LLMs lack diversity and increasing the sampling temperature can degrade text quality before sufficient diversity is achieved. Existing decoding methods address this trade-off through token-level filtering. While they provide heuristics about how valid tokens appear in the next-token distribution, their measurements happen at the sequence level. This leaves a gap between where the methods are actually proposed and where they are being evaluated. We introduce a framework for diagnosing and quantifying the validity-diversity trade-off directly at the token level. Our approach combines tasks with verifiable output spaces and an automated procedure for open-ended tasks. Using our framework, we identify two failures in LLMs: First, order calibration: valid tokens are not reliably ranked above invalid tokens. No decoding method based on top-token filtering can recover enough valid tokens without including many invalid tokens. Second, shape calibration: probability mass is overly concentrated only on a few valid tokens while having a heavy tail of mixed valid and invalid tokens, so maintaining high diversity will inevitably expose the invalid tokens in the tail. We formalize both mechanisms and show that local failures compound across decoding steps, producing strong sequence-level losses in diversity. Across 14 language models spanning multiple families and scales, we find that diversity collapse is not merely a limitation of particular sampling heuristics, but a consequence of order and shape miscalibration in the LLM distribution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.