Cooperative Words, Adversarial Actions: Action-Level Ambiguity in Embodied Multi-Agent Systems
Abstract
Embodied multi-agent systems increasingly rely on foundation-model agents to coordinate long-horizon physical tasks, making reliable attribution of team failures a growing safety challenge. We identify action-level ambiguity as an overlooked vulnerability, where a compromised agent can remain locally plausible while systematically reducing collective progress. We introduce a stealthy action-level attack that induces contextually defensible but lower-contribution actions while preserving cooperative communication and faithful observation reports. Across diverse embodied benchmarks, coordination frameworks, and agent backbones, the attack substantially degrades team performance while remaining difficult to attribute, achieving an overall stealth rate of 0.9980 across three investigator models, three auditing strategies, and 8829 episode-level audits. In contrast, overt, lazy, and randomly suboptimal agents are substantially easier to identify. These results expose a gap between behavioral plausibility and contribution integrity. To reduce this attribution gap, we further develop a contribution-aware, log-based attribution method that reasons about each agent’s marginal task contribution and interference with teammates’ intents, substantially improving compromised-agent identification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.