Localization is not Attribution: Diagnosing Who-and-When Blame in Multi-Agent LLM Systems
Abstract
When a multi-agent LLM system fails, two questions must be answered to enable repair: when did the trajectory go wrong, and who is responsible. We show that these two questions are far from equally hard. Across four frontier LLM judges on the Who&When benchmark of real multi-agent failures, every method we test—including transcript-reading baselines, structured step-finding, and a strong leave-one-out evidence attribution—localizes the failing step far more often than it correctly attributes the responsible agent. We quantify this gap with a diagnostic metric, rsw: the fraction of correctly localized cases in which the blamed agent is nonetheless wrong. rsw stays stubbornly high (0.19–0.24) for all methods and, crucially, rises on the annotation subsets that multiple judges agree are most trustworthy—evidence that rsw reflects a genuine model limitation, not label noise. Diagnosing the failure, we trace the error to role-charter collapse: when two agents in a real trajectory play overlapping roles, judges summarize them into indistinguishable profiles and then misattribute blame. We propose Role-Op Attribution (ROA), which builds distinctive role charters via leave-one-out contrast and aligns them to per-agent operation signatures. ROA is the first method to significantly reduce rsw (-0.036; 95% CI [-0.068,-0.003], paired bootstrap, n=632) while nudging joint who+when accuracy upward (+0.027; 95% CI [-0.003,0.059]), at 40% lower judge-input cost than the strongest baseline. Component ablation shows the two mechanisms are coupled rather than independent tricks, and self-consistency voting fails to close the gap, confirming that attribution demands a structural fix rather than more sampling. Finally, the same localization–attribution dissociation generalizes to a vision–language multi-agent setting: perception-layer faults are systematically mis-blamed on the reasoner across five visual modalities, both under controlled injection and on 665 naturally occurring VLM failures (a 0.36 who-attribution gap whose 95% CI excludes zero). We release all code and per-case verdicts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.