acceptodds
Under review as a conference paper at ICLR 2027

SIMULATING RESPONDENTS, NOT SINGLE QUESTIONS: COHERENT SURVEY GENERATION WITH LARGE LANGUAGE MODELS

Abstract

Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real-world questionnaires, however, typically require each respondent to answer a sequence of related questions. A simulated respondent should therefore exhibit coherent preferences across the entire questionnaire, rather than merely produce accurate distributions for isolated items. Existing single-item simulation methods can closely match item-level response distributions, but they do not accurately reproduce how the same respondent answers a complete survey. To address this limitation, we propose FullRespondent-LLM (FR-LLM), a framework for simulating complete virtual survey respondents. FR-LLM first fine-tunes two specialized LLMs: a marginal model that estimates the response distribution of each item and a respondent-level autoregressive model that captures dependencies among answers across the questionnaire. We combine these models through Marginal-Constrained Joint Projection (MCJP), which projects the autoregressive joint distribution onto the set of distributions satisfying the item-level marginals learned by the marginal model. By modeling item distributions and cross-item relationships separately, FR-LLM generates complete questionnaires that reproduce realistic cross-item relationships while retaining the item-level accuracy of strong single-item simulators. Experiments on two real-world social survey datasets show that FR-LLM more accurately reproduces response patterns across multiple questions, while maintaining competitive single-item accuracy and generalizing better to unseen respondent populations and survey questions. We also conduct a small commercial survey, use the responses generated by each method to make the same business decision, and compare the resulting profits. FR-LLM produces the highest profit, demonstrating its potential for practical survey-based decision making.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.