ReLoC: Recursive Local Causal Modeling for Explaining and Predicting LLM-Based Multi-Agent Behavior
Abstract
Large language model (LLM)-based multi-agent systems (MAS) are increasingly deployed to solve complex tasks; thus learning how a local intervention on one agent causally affects the overall outcome is necessary. As the MAS scales, however, the space of joint interventions grows combinatorially, making a global causal model prohibitively expensive and preventing measured effects from extending to unseen intervention queries. We introduce Recursive Local Causal Modeling (ReLoC), which progressively builds an approximate global causal model by learning and composing local causal models of individual agents, each built from a few interventions. Specifically, ReLoC learns how each agent responds to interventions on task-defined concepts. Starting from the outcome agent, ReLoC recursively learns how upstream agents produce these concepts, following the communication graph back to the task inputs. ReLoC eventually composes these local models into the global model. Across three benchmarks, three topologies, and four LLMs, ReLoC accurately predicts how each agent responds to concept interventions, how these changes propagate through the system, and how the final answer changes under unseen interventions, consistently outperforming an LLM that directly reasons over the full MAS under the same intervention budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.