acceptodds
Under review as a conference paper at ICLR 2027

Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas

Abstract

Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents, to reduce the costs of survey research, market analysis, and policy experimentation. The validity of LLM-based synthetic personas as substitutes for particular respondents depends on whether they reproduce individuals' decisions. We examine what information helps synthetic respondents better predict each individual's later choices. We design experiments that evaluate synthetic respondents constructed with progressively richer information about each individual, beginning with no personal information, then adding demographics, personality traits, cognitive scores, and finally the respondents' earlier survey choices as behavioral history, yielding five conditions. We use a two-wave panel of 845 US adults who completed measures of 14 behavioral biases (spanning risk, time preferences, overconfidence, and reasoning), so each respondent's earlier answers provide a human test–retest benchmark. When earlier answers serve as behavioral history, all items that score the target bias are withheld. At the population level, the average number of biases per respondent in every condition is close to the human average (7.1–8.1 biases, against 7.1 for humans). This aggregate similarity, however, masks substantial differences in variance structure: persona descriptions recover only 53–67% of the amount of human between-person variation, whereas adding behavioral history restores the overall amount of variation to approximately the human level. At the individual level, predictive agreement is much weaker: description-based personas achieve only 7–12% of the informedness observed in human test–retest responses, while adding behavioral history raises this to 28%. Across demographic groups, the cumulative condition including behavioral history has the highest estimated informedness in all 17 groups, whereas description-based conditions provide little or no information for some groups. We also find that synthetic responses exhibit stronger education- and income-related differences than human responses. For LLM synthetic personas, a respondent's past answers add more to individual-level prediction than a description of who they are.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.