Concordia Island: A Sandbox World for Evaluating the Psychology of Generative Agents
Abstract
Generative agents have demonstrated promising fidelity in simulating human behavior on lab experiments and surveys. But can such agents sustain coherent lives and psychological responses across an entire simulated society with real-world complexity? We introduce Concordia Island, a scalable, open-ended sandbox world in which 100+ agents autonomously live, work, socialize, and adapt to significant life events. To systematically evaluate these simulations, we present a comprehensive evaluation suite for synthetic psychology. This suite integrates validated psychometric instruments for measuring emotion, personality, and life satisfaction with open-ended reflections about meaning and identity to audit the validity of agents' constructed inner experiences. We organize our evaluation around three classes of metrics: 1. longitudinal tracking of agent baselines and their responses to exogenous shocks such as job displacement, 2. psychometric reliability and validity and 3. qualitative analysis via an automated ethnographer and qualitative coding pipelines for systematically evaluating the large simulation artifacts that are produced. Across a model comparison spanning agent decision-logics and base LLMs, we demonstrate that agent self-reports exhibit reliable factor structure, context-appropriate emotional responses to life events, and meaningful individual differences—while also revealing systematic failure modes such as unrealistic attractor basins in long-run simulations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.