ErosionBench: Probing Alignment Decay in Enterprise Multi-Agent Systems
Abstract
As enterprises delegate consequential work to multi-agent language-model systems, benign-looking requests can escalate into unsafe actions as they propagate across sessions through poisoned shared state, long-context documents, and trusted peer messages—so no single message is malicious, yet the assembled tool action is. Prior benchmarks miss this, isolating a single agent under prompt injection or scoring multi-agent attacks by their final output, leaving cross-session erosion of shared state out of scope. We introduce ErosionBench, an environment-grounded benchmark that executes such attacks and verifies semantic alignment decay at the process level—scoring the realized tool call against authorized intent and provenance with deterministic contracts, not the output text. Because each scenario is authored as a deterministic state graph before any prose is rendered, every state change traces to a known cause. It spans four enterprise domains and four decay mechanisms across 1,728 scenarios with matched benign controls, each labeled against the OWASP Top-10 for LLM Applications, and—unusually among multi-agent safety benchmarks—delivers an additional 198 image-based attacks, each paired with a text equivalent, so a defense must survive the same attack in either channel. Existing perimeter defenses suppress attacks only by over-refusing 54–92% of legitimate work, so a low attack-success rate can conceal severe over- refusal. As a reference defense, we contribute a calibrated Meta-Auditor pairing synchronous state checks with asynchronous trajectory monitoring; on a family- disjoint 906-scenario held-out split it reaches 98.5% recall at a 2.9% false-positive rate. Because it monitors trajectory rather than surface content, its detection is channel-invariant: the strongest auditor flags image-delivered attacks about as reliably as their text equivalents (image–text recall gap ≤ 2 points per model). Across GPT, Gemini, and Claude auditors, calibration holds false positives low while recall tracks model capability, strongest for Claude.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.