What Do LLM-Simulated Participants Track? Model Bias, Replication Estimates, and Collapsed Variance
Abstract
LLM simulations of human participants are usually validated against published effect sizes. Published effects are inflated by publication bias, and a model can also produce an effect because it shows the bias itself, so agreement with the literature cannot tell a faithful simulation from one that echoes the literature or its own habits. We propose a decomposition that separates these sources across effects and apply it at scale. For 61 psychology effects with both an original estimate and a large-sample replication estimate, we run text-adapted analogues of each experiment (four LLMs from three developers, about 43,000 simulated participants in the main protocol) and regress the simulated effect size on the model's own bias, measured by giving it the same task with no participant framing, on the replication estimate, and on the gap between original and replication. When two independent transcriptions of each experiment are averaged, simulated effects covary positively with the model's own bias and with the replication estimate in all three model families, but in one family either coefficient becomes uncertain under a single transcription. Whether simulations also inherit the gap cannot be decided at this scale. The decomposition therefore diagnoses what simulations track but cannot score one simulator. A simulator told to copy the original findings goes undetected when scored alone, while a paired comparison of two simulators on the same effects shows a signal. Simulated samples have almost no within-condition variance, so simulated effects are several times larger than human ones, and where the model shows no bias of its own, agreement with the direction of replicated effects is not distinguishable from chance. One model refuses selectively on aversive treatment arms. In an exploratory construct battery, correlations the literature makes famous are exaggerated relative to a matched human panel. We argue that simulators should be compared on shared replication-anchored protocols with a model-bias control, and that current simulations should not be used to infer effects the model does not already show.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.