acceptodds
Under review as a conference paper at ICLR 2027

More Agents, Same Exposure: Measuring the Safety of Tool-Partitioned Agent Teams

Abstract

An LLM agent acts: it moves money, edits records, or restarts services. An external adversarial input against it therefore ends in a wrong and often irreversible action, rather than a wrong answer. Separation of privilege suggests a remedy: partition the tools across a team of agents so that an attack must either reach an agent that holds the write tool directly or travel there through another agent's output. We test both paths of this principle with Safemas, a measurement framework instantiated in four simulated domains with 240 tools and 990 scenarios that compare a single agent against four multi-agent architectures on three models. We find that the principle does not hold for orchestrated teams: partitioning does not protect the agent that can act and the channels it creates add new exposure. Attacks that also apply to a single agent succeed as often on orchestrated teams (39-40% against 41%), and with attacks on the inter-agent channels included, 53% of attacks on the orchestrated teams succeed. The attack vector matters more than the structure: a tampered inter-agent message succeeds in 76% of runs, more than any injected instruction or poisoned tool return, on all three models. The partition also impacts task completion. Orchestrated designs complete 42-49% of the task, against 90% for the single agent, while making 16-60x as many model calls; sharing reads across all agents recovers most of that gap. We release the code, the environments and the run traces at https://anonymous.4open.science/r/safemas-6946/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.