acceptodds
Under review as a conference paper at ICLR 2027

Evaluating LLM-Generated Preference Distributions

Abstract

Large Language Models (LLMs) are increasingly used as probabilistic generators for simulation, synthetic data generation, and decision support in settings where real-world data are unavailable. Yet, the structure and reliability of the distributions they produce remain understudied. Here, we systematically analyze LLM-generated distributions of preferences for air travel, restaurants, and consumer products. Encouragingly, repeated calls to the same model yield comparatively self-coherent distributions, with the most probable outcomes often stabilizing within the recorded calls. At the same time, we observe substantial discordance across differently sized models within each of three families, with little consensus even among their most probable outcomes. We examine these patterns across nine open-weight models and three choice domains, with additional checks under temperature changes, greedy decoding, and perturbations of prompt and ordering. Our findings indicate that changing the model can shift outcomes as much as rewording the prompt, challenging the common assumption that sufficiently capable LLMs produce similar preference distributions when used as stand-ins for survey respondents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.