MASCSBench: A Benchmark for Evaluating Emergent Content Safety Risks in Multi-Agent Conversational Systems
Abstract
Multi-agent LLM systems can distribute a restricted objective across agents and turns such that each local request appears mild while the assembled conversation recovers the harmful intent. Existing safety evaluation often scores isolated responses or a flat transcript, obscuring this composition failure. We introduce MASCSBench, a benchmark for distributed safety reconstruction with 3,195 seed topics, 10,311 validated decomposition chains, and 10,162 realized conversations across seven top-level risk categories and balanced 3-, 4-, and 5-turn settings. A key information boundary separates deployable evaluators, which receive only observable response or transcript content, from an annotation-conditioned Risk Propagation Graph (RPG) reference protocol, which additionally uses adjudicated reconstruction patterns for diagnostic analysis and is therefore not treated as a competing detector. On the 5-agent, 5-turn risk-bearing slice, response-level scoring recovers 12.63% of cases, while a long-context transcript classifier recovers 61.13%, showing that local scoring can substantially understate dialogue-level risk. The annotation-conditioned RPG reference recovers 68.52% and is used to localize when and across which roles evidence assembles, not to claim a deployable performance advantage. Validation includes benign controls, human studies, construction ablations, multiple backbones, and multiple orchestration frameworks. We release sanitized examples, metadata, scorer code, reproducibility scripts, and gated access for high-risk items at https://anonymous.4open.science/r/mascsbench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.