Evaluating Textual Diversity Metrics with Controlled Mixtures of Diverse Perspectives
Abstract
The prevalence of large language models (LLMs) as tools for information access and creativity necessitates the development of rigorous methods to assess the diversity of LLM-generated content in open-ended contexts. While many diversity metrics have been proposed in the literature, incorporating both varied representations of text and methods for measuring diversity, we identify a significant gap between existing abstract studies of diversity, and the practical realities of measuring diversity in text. Drawing on the interdisciplinary framework proposed by Stirling, which decomposes diversity into three necessary dimensions (variety, balance, and disparity), we add a fourth requirement specific to open-ended natural language, invariance to composition, and operationalize the framework to develop a generalizable procedure for validating proposed diversity measurements of natural language grounded in real datasets of diverse text. Concretely, we: (1) Demonstrate how to generate controlled mixtures of diverse real-world, open-ended text from existing datasets with processes that make the four fundamental diversity criteria directly accessible; and (2) Empirically measure and statistically test which diversity metrics correspond with the underlying conceptual diversity. Our evaluation bridges the gap between theoretical notions of diversity and empirical work in natural language processing, providing grounding for diversity measurement in open-ended text and a methodology for validating future proposed metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.