Do Collaboration Metrics Track Multi-Agent Success? A Cross-Benchmark Measurement Audit
Abstract
Multi-agent LLM systems are increasingly deployed on complex, long-horizon tasks, but are usually evaluated through final outcomes or structural scale, neither of which tells us whether the collaboration itself was effective. Successful trajectories can contain poor coordination, and failed trajectories can still involve useful information exchange. We therefore investigate whether trajectory-level collaboration metrics provide stable signals of task success across domains, models, and system designs. We evaluate five metrics of collaboration quality, namely a holistic judged communication score, judged redundancy, judged downstream usage, judged contradiction rate, and semantic evolutionary distance, across four multi-agent benchmarks spanning algorithmic coordination, software engineering, and enterprise workflows. The study covers five language models and 18 model-benchmark configurations. Four of the five metrics fail to hold up. The narrower per-event judged measures survive multiplicity correction in at most six of 18 configurations, while evolutionary distance changes sign across settings and even within a single benchmark, so neither can be read as a directional quality score. The fifth, a holistic communication score, is associated with task success in 14 of 18 configurations (–), every one surviving correction, but its rubric partly references task progress. We next ask whether the same signal transfers to a different outcome. In a secondary analysis of 1,500 trajectories from three privacy-constrained benchmarks, no general collaboration metric survives multiplicity correction as a predictor of information exposure, and this non-detection holds whether exposure is measured across the whole trajectory or restricted to failures that never reach the final output, so the privacy analysis can rule out large associations while smaller effects may still have gone undetected. Under this instrument, the measures that are cleanly separable from the task-success outcome carry little signal that transfers across benchmarks, and the one measure that does carry signal is not cleanly separable from it. Collaboration metrics should therefore be validated for a specified outcome and setting rather than treated as general indicators of collaboration quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.