From Simulation to Evidence-Grounded Decision Analysis: A Benchmark and Decision-Support Framework for Multi-Agent Social Systems
Abstract
Most LLM-based social-simulation studies evaluate realism, social intelligence, or task success within a simulated world. We ask two related questions: whether agents exhibit adaptive social strategies beyond role-consistent responses, and how execution traces and simulation results can support disciplined decision analysis. We present , an executable benchmark with social-media and commercial-negotiation environments such as role profiles, private goals, visibility constraints, and deterministic state transitions. We evaluate agents with rule-based outcomes and five LLM-as-a-Judge dimensions, while also analyzing the strategies they employ, finding that even the excellent LLMs only perform simple strategies without applying complex methods. We further study a simulation-to-decision workflow that transforms traces into conditional reports with triggers, alternatives, trade-offs, and exit conditions. Our report ablations show that trajectory summaries are more useful than equally budgeted raw traces. Matched simulation evidence can provide additional decision-relevant content, but current results do not establish that simulation-grounded reports consistently outperform carefully structured scenario-only reports. Accordingly, evaluates trace-grounded report properties rather than real-world decision outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.