S2Bench: A Survey-Grounded Benchmark for LLM-Based Social Simulation
Abstract
LLM-driven social simulations enable controlled and repeatable experiments that would be costly, slow, or unethical with real people. Evaluating whether these simulations faithfully reflect human behavior remains challenging; existing work largely relies on real-world survey responses as quantifiable records of human choices. Yet current survey-based benchmarks have three limitations: they evaluate only static settings with fixed respondents and questions, lack a unified evaluation protocol, and remain narrow in coverage and scale. To address these limitations, we introduce S²Bench, a large-scale benchmark for survey response prediction. It defines four settings by crossing seen and unseen respondents with seen and unseen questions. We also provide a public evaluation platform that scores submitted methods under identical conditions and ranks them on a live leaderboard, and release the benchmark construction pipeline. Applied to 15 real-world survey datasets across 234 waves, the pipeline yields nearly 0.4 billion respondent–question–answer triples. We evaluate prompting- and training-based methods across 13 LLMs and find that even the strongest models reach only 40–55% accuracy. Although scaling and reasoning modes improve performance, the gains diminish quickly, suggesting that the ability to simulate a person has improved far more slowly than general abilities such as knowledge and reasoning. Fine-tuning transfers well to new respondents, but much less to new questions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.