MOSAIC: Unified Outcome-Aware Security Auditing for LLM-Based Multi-Agent Systems
Abstract
Large Language Model (LLM)-based multi-agent systems (MAS) enable complex tasks through communication and collaboration among specialized agents, but their interconnected nature introduces security risks beyond the scope of conventional LLM and single-agent safeguards. Existing MAS defenses focus on complementary aspects such as access control, anomaly detection, or task-failure prediction, yet lack a unified mechanism for determining whether an attack actually compromises the system, where its effect occurs, and what evidence supports the decision. We propose **MOSAIC**, a unified outcome-aware security auditing framework that integrates attack detection, realized-outcome assessment, structural localization, and evidence-grounded attribution for MAS. We further introduce **MOSAIC-Bench**, a systematic benchmark that distinguishes attack exposure from realized compromise, and **MOSAIC-Auditor**, a graph-grounded auditor for reasoning over multi-agent execution trajectories. Experiments demonstrate the effectiveness of MOSAIC for comprehensive MAS security auditing and highlight the necessity of outcome-aware, structurally grounded reasoning beyond existing safety paradigms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.