Interpreting Emergent Extreme Events in Multi-Agent Systems
Abstract
Large language model-powered multi-agent systems are powerful tools for simulating complex human-like systems. Their interactions often produce extreme events whose origins remain obscured by the black box of emergence. Interpreting these events is critical for system safety. To our knowledge, this paper proposes the first unified framework for retrospectively interpreting emergent extreme-event risk across time, agents, and behaviors in LLM-powered social simulations, answering three questions: When does the event originate? Who drives it? And what behaviors contribute to it? Specifically, we adapt the causal-order-constrained Shapley value to attribute reference-relative extreme-event risk to each agent action at each time step, i.e., scoring the action by its marginal contribution in replay. We then aggregate the attribution scores over time, agents, and behaviors to quantify the risk contribution of each dimension. Finally, we design metrics from these contribution scores to characterize extreme events. Experiments across diverse multi-agent system scenarios (economic, financial, and social) demonstrate how our framework characterizes extreme events through their temporal, agent, and behavioral contributions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.