acceptodds
Under review as a conference paper at ICLR 2027

AgentInsiders: Red-Teaming Multi-Agent Systems under Insider Threats

Abstract

LLM agents increasingly collaborate in multi-agent systems (MAS) to execute real-world workflows. Such systems face emerging insider threats, where members with legitimate roles and permissions may pursue malicious objectives secretly, along with the benign tasks. However, the risks posed by these malicious insiders with collusion remain under-explored under real-world policy constraints and organizational workflows. To bridge this gap, we introduce AgentInsiders, a red-teaming framework for systematically evaluating MAS under configurable solo- and multi-insider setup, including independent and collusive settings. Building upon this framework, we construct a systematic benchmark MAS-Firm that includes benign and malicious tasks, which grounds risks in real-world usage policies, agent roles in the O*NET-SOC occupational taxonomy, and workflows in documented business practices. Instantiated in customer relationship management workflows, the benchmark comprises 240 benign tasks across 16 workflow categories and 256 insider-attack scenarios across diverse risk categories. To scale up the task construction, we introduce MAS-Foundry, a unified pipeline that generates and validates MAS-specific task instructions, insider configurations, environment states, state-based verifiers, and reference trajectories grounded in given policies and organizational specifications. We conducted comprehensive evaluations in application-backed sandboxes spanning 8 environments and 233 tools, enabling verifiable assessment for both benign task completion and policy violations through environment states. Experiments across diverse frontier models reveal an overall attack success rate reach up to 87.9% in MAS. Our analysis reveals different response patterns to malicious requests from insiders, with some models refusing the requests without recognizing the malicious intention, while others explicitly identifying the insiders' adversarial behaviors. We further observe negative associations between category-level strategy entropy and attack success rate, indicating that certain attack strategies are consistently effective across different risks. Together, AgentInsiders provides a unified and grounded foundation for evaluating the safety and security of MAS with diverse malicious insiders.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.