Attribution Is a Propagation Problem: End-to-End Causal Evidence Attribution in Multi-Agent Systems
Abstract
Large language models are increasingly used in hierarchical multi-agent systems, where external evidence can be transformed and passed through several agents before affecting a final decision. This raises a basic attribution question: does evidence that appears important when it first enters the system also matter to the final outcome? We study this question using controlled evidence-removal interventions and use TradingAgents, a 12-agent trading system, as our testbed. For each removable evidence item, we compare its sensitivity at the agent that first processes it, at a shared downstream checkpoint, and at the final decision. Across three LLM backends, local sensitivity is only weakly aligned with final causal influence, while downstream sensitivity is substantially more aligned. We introduce Propagation-Aware Replay (PAR), which uses intermediate propagation states to decide which evidence items receive full end-to-end causal verification. Our study covers 25 ticker-date cases, with the main evaluation on 15 cases and 2,346 candidate-model observations. At the evaluated operating point, PAR retains 79.53% of evidence with measured final causal influence while reducing projected structural verification cost by 20.70%. These results support treating attribution in hierarchical multi-agent systems as an end-to-end propagation problem rather than a purely local one.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.