acceptodds
Under review as a conference paper at ICLR 2027

Hallucinated Heterogeneity: LLM Social Simulations Exaggerate Who Responds to Interventions

Abstract

Large language model (LLM) social simulations are increasingly used as a fast, cheap substitute for human experiments, and are typically validated on whether they recover average treatment effects. But targeting and personalization depend on a different question: who responds? We introduce H-Bench (Heterogeneity Bench), a benchmark of 138 randomized experiments spanning economics, politics, health, and other domains, which uses causal inference to test whether simulations recover treatment effect heterogeneity, from subgroup differences to the value of targeting decisions, against randomized human outcomes. Across nine LLMs, we find Hallucinated Heterogeneity: simulations track average effects far better than who responds, and they exaggerate how much people differ. Claude Sonnet 5's average effects correlate with human estimates at r=0.74, but its subgroup differences in treatment effects correlate at only r=0.14 (p<0.001 for the difference), a gap that persists after correcting for human sampling noise (0.80 vs. 0.29, p=0.04). Eight of nine models overstate heterogeneity by 1.4–2.7×, and LLM-based targeting barely beats random assignment, even after calibration on human data. We explain why this failure arises and goes undetected. Simulated responses are often far from human ones, but the errors are similar with and without treatment and cancel in the average effect, so standard validation passes. And the heterogeneity models produce follows familiar divides such as partisanship, where models exaggerate it without identifying which groups respond more. Together, these results show that LLM social simulations must be validated at the level at which they will be used before they can support targeting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.