Risk Lies in the Relations: Relation-Aware Collaboration Graph Defense for LLM Multi-Agent Systems
Abstract
Large language model (LLM) multi-agent systems are commonly defended by classifying messages or auditing agents in isolation. Safe components, however, do not guarantee safe collaboration: individually benign messages, retrieved documents, tool responses, and memory records can jointly trigger harmful actions after transformation, delegation, and downstream authority amplification. We introduce RACG, a relation-aware collaboration graph framework that treats multi-agent defense as intervention over an asynchronous execution process rather than anomaly detection over isolated text. RACG represents execution dependencies in a typed event-level partial order and learns the safety and utility effects of typed interventions from matched checkpoint replays. At deployment, affected-cone modeling and uncertainty-constrained search select executable intervention sets and regenerate the affected continuation from an action-valid rollback cut. Across nine benchmarks, RACG achieves 13.81% macro attack success rate and 70.32% defended task success, compared with 15.54% and 69.31% for Full-Graph Late Fusion. For joint interventions, it reduces pair/triple risk MAE from 0.104/0.139 to 0.086/0.113, while eMEIS attains 0.071 regret by evaluating 27.54% of candidate sets. These results support relation-aware intervention as a practical basis for secure multi-agent collaboration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.