When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
Abstract
Large language models (LLMs) are increasingly used as synthetic users: stand-ins for human respondents whose simulated answers inform product, policy, and market decisions. We ask when this substitution is valid and when it fails. Using a unified evaluation protocol, we study four models spanning two model families and an 8B-to-frontier capability range across two independent domains of real human responses: U.S. social attitudes from the General Social Survey (GSS) and cross-cultural values from the World Values Survey (WVS). We benchmark each model against non-LLM baselines fit on held-out human data. Across both domains and model families, we find two consistent failures under the demographic prompting and survey-simulation protocols we test. First, at the individual level, no LLM outperforms the strongest non-LLM baseline. Under proper scoring rules (log-loss and Brier score), every model assigns less probability to the observed human response than a demographic baseline; this gap persists under distance-aware scoring, and on WVS, every model also falls well below the baseline in exact-match accuracy. Second, the models over-determine demographics: they treat demographic identity as more predictive of attitudes than it is among real human responses, most clearly for the Claude models and Llama-70B, though our small model sample overstates this gap. In the two within-family size comparisons we test, the larger model remedies neither failure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.