acceptodds
Under review as a conference paper at ICLR 2027

SCE-Bench: Multi-scale Behavioral Alignment of LLM Agents in Real-world Microeconomics

Abstract

The adoption of Large Language Models within Agent-Based Modeling enables the simulation of economic behaviors, yet the empirical validity of these LLM-driven agents remains largely unverified. Current benchmarks often emphasize normative optimization or game-theoretic tasks, omitting the temporal consistency and individual heterogeneity necessary for realistic economic modeling. In response, we introduce SCE-Bench, derived from the Survey of Consumer Expectations using over 70,000 nationally representative observations from 2016 to 2023. We establish a hierarchical evaluation protocol that measures alignment at both the aggregate level, through temporal correlation and distribution matching, and the individual level, using a Monte Carlo-based approach to assess individual differences and statistical indistinguishability. Our empirical analysis covers comparisons across model characteristics, prompt variations, and temporal contexts. The results indicate that LLMs often exhibit an “Average Person” bias and lack precise temporal tracking. Meanwhile, cost-effective open-weight models emerge as viable alternatives for large-scale simulations. These findings, along with insights into feature selection, verbalization, role-playing perspective, and temporal contexts, establish practical guidelines for configuring LLM economic agents in realistic simulations. Finally, we further validate the generalization of our framework and findings on cross-domain tasks and downstream simulation systems. Dataset, code, and experimental results are available at https://anonymous.4open.science/r/SCE-Bench-BBA7.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.