High Scores, Narrow Vocabularies: Measuring Cross-Trial Diversity in Large Language Models
Abstract
Large language models are increasingly used in creative industries as sources of novel ideas, treating output variation as evidence of genuine creative exploration. The Divergent Association Task (DAT), a validated creativity benchmark, has repeatedly produced at- or above-human level scores for LLMs, but a single-trial score cannot distinguish broad semantic exploration from repeated draws on a narrow, high-scoring vocabulary. We demonstrate this dissociation directly: LLMs can attain human-level or above-human DAT scores while showing substantially narrower cross-trial vocabularies than humans. We test five models, three open-source and two frontier, with 100 trials per condition under varied prompts, reasoning instructions, and decoding parameters, assessed using five metrics: DAT score, word overlap, vocabulary breadth, entropy, and semantic-region overlap. Prompt wording and reasoning instructions are the strongest predictors of DAT score, but they do not reliably increase lexical diversity. By contrast, widening the sampling pool reliably increases diversity without a corresponding gain in score for most models. Therefore, neither lever, within the standard decoding settings, closes the gap between benchmark performance and genuine exploration. Frontier models show the most extreme version of convergence, showing the highest cross-trial overlap and the narrowest vocabulary of any model tested. These findings show that human-like DAT performance does not imply human-like diversity. We propose multiple metrics of cross-trial diversity as a behavioral diagnostic of generative convergence that single-trial scores cannot reveal. As LLMs become routine in creative work, such convergence may homogenize otherwise independent outputs, warranting caution in treating the properties of individual outputs as evidence of genuine creative diversity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.