Rethinking the Evaluation and Optimization of LLM-based Social Simulation
Abstract
LLM-based social simulation is a promising complement to traditional social science methods such as surveys and behavioral experiments. A core question in this area is how to evaluate the fidelity of LLM-simulated human behavior and, in turn, how to optimize LLMs toward it. Prevailing practice evaluates by accuracy, checking whether the model selects the single response observed from a human, and accordingly trains the LLM to reproduce this one hard label. However, human behavior is inherently subjective: the same person in the same situation may reasonably act in different ways, so an observed response is only one draw from an underlying response distribution, which renders accuracy-based evaluation unreliable and hard-label training misleading. To address these problems, in this paper, we first introduce the subjectivity coefficient, an entropy-based quantity that distinguishes objective tasks such as coding from subjective tasks such as social simulation, and then use it to systematically analyze how accuracy-based evaluation and hard-label training fail as subjectivity grows. Based on the subjectivity coefficient, we further propose Subjectivity-Adaptive soft-Label Training (SALT): it pools observed outputs from semantically nearby inputs into soft distributional labels, with an aggregation radius adapted to the estimated subjectivity of each input; in the near-objective limit the neighborhood shrinks, so SALT naturally falls back to standard single-label training. Moreover, since existing datasets record only single observed responses and thus cannot support distributional evaluation, we construct SUBJSIM, a benchmark of 19,300 contexts covering 193 annotators and 100 subjective questions. Since real-world data typically provide only a single observation per input, our experiments train models from single observed outputs while evaluating them against the full response distributions, thereby verifying the feasibility of our method in realistic settings. Extensive results on SUBJSIM demonstrate the advantages of our method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.