Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
Abstract
Non-verbalised or secret collusion may impose a ceiling on transcript-level monitoring in multi-agent deployments. Internal model activations are known to provide an additional monitoring surface for single-agent phenomena such as deception, but detecting collusion within an unknown subgroup is ultimately a joint inference task over multiple agents. This paper derives collusion detection statistics based on the composite hypothesis testing literature. Using per-agent activation probes as sensors, we then evaluate these principled statistics alongside heuristics on NARCBench, a novel benchmark for evaluating multi-agent collusion detection under environment distribution shift. Across four open-weight models, both our group-level statistics and heuristics transfer zero-shot to NARCBench's holdout environments. Notably, the best method on every model reaches 0.99 to 1.00 AUROC when transferred from committee deliberations to a steganographic collusion task that unprimed transcript monitors fail to separate. We then evaluate detector transfer to third-party environments with emergent, rather than instructed, collusion. In one of these environments, only 2 of 229 collusive episodes verbalise the coordination in the transcript. We find evidence of partial transfer to these environments, indicating that probes fitted to instructed collusion may not fully capture emergent collusion signals. In line with our theoretical analysis, we find that no single statistic or heuristic dominates empirically across tasks. Overall, we establish multi-agent interpretability as a complementary monitoring technique, while highlighting the need to evaluate detector performance on both emergent and instructed collusion risks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.