When Does Coordination Induce Correlated Failure in Multi-Agent LLM Reasoning?
Abstract
Multi-agent LLM systems increasingly rely on a meta-agent that issues a shared, explicit plan to coordinate otherwise independent executor agents, but whether such coordination improves or harms reasoning reliability is not well characterized. We present Controlled Deliberative Decentralized Execution, a minimal framework that isolates coordination strength as a single manipulable variable. We evaluate seven open-weight models across three reasoning domains. We introduce a two-statistic diagnostic framework—a correctness intraclass correlation () and an excess pairwise agreement statistic ()—designed to separate genuine correlated failure from benign narrowing of each executor's individual answer distribution. Strong coordination reduces accuracy relative to self-consistency in 19 of 21 (model, domain) cells, and this directional harm is unanimous at the model level: all 7 models are harmed on at least two of three domains. Critically, the underlying mechanism is not uniform: only three of seven models exhibit the correlated-cascade signature (rising ) while the remaining four show comparable harm through a distinct, non-correlational route. This mechanistic split tracks neither model scale nor family. One robust exception survives across independent measurement corrections: Qwen2.5-3B on MATH shows a significant accuracy improvement under coordination, confirming that agreement's decoupling from correctness cuts in both directions. We argue that executor agreement is an unreliable proxy for reliability and that single-mechanism governance heuristics are miscalibrated for a meaningful fraction of deployed systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.