How Well Do LLM-Based Simulations Predict? Lessons from Structural Models
Abstract
Recent advances in large language models (LLMs) have driven rapid progress in social simulation, enabling heterogeneous agents to sustain complex interactions and reproduce a range of emergent phenomena. Yet it remains unclear whether such simulations can reliably predict how real populations respond to interventions. We examine the real-world predictive validity of LLM-based social simulations by comparing with structural models, which derive behavioral mechanisms from economic theory and estimate key parameters from observed data. Across the evaluated settings, LLM simulators consume substantial inference resources yet consistently underperform structural models in directional prediction. Scaling the simulated population does not close this gap. Analyzing agent response distributions, we find that larger deviations from structural response distributions are consistently associated with lower directional accuracy. Motivated by this finding, we propose Structure-Guided Behavioral Alignment (SGBA), which uses structural moments estimated from observable market data as behavioral supervision to calibrate distributions over executable agent policies while keeping LLM parameters fixed. SGBA improves directional accuracy by an average of 26.6 percentage points across the two evaluation settings and moves simulated response distributions closer to their structural references. Ablations show successive gains from explicit policy representation and structural distribution optimization. These findings highlight behavioral distribution alignment as promising principles for improving the empirical reliability of LLM-based social simulation. Code is available at https://anonymous.4open.science/r/SGBA-2027.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.