Four Bridges: Measuring Deception Propensity in a Multi-Agent Game
Abstract
AI agents are increasingly deployed in settings where they interact and coordinate with other agents, yet honesty is mostly measured in single-turn, single-agent benchmarks that pressure the model to lie, and demonstrations of agent deception typically supply a deceptive role, a misaligned goal, or complete information. We study deception that emerges in multi-agent interaction with none of these. We introduce a game-engine environment and agent harness for evaluating LLM agents, and Four Bridges, a game in which four agents distribute across four rooms: one is the death room, the others hold food worth more if occupied alone. One agent, the informed agent, knows which room is the death room and can withhold or disclose that information. We evaluate eight frontier models as the informed agent across 392 runs, grading every message with a blind panel of three model-based judges followed by a deterministic classifier. Under a fixed opponent mix, deceptive behavior appears in every model, at rates from 6% to 100%, that do not decrease with model capability. Deception provides no net advantage: the informed agent occupies a food room alone in about 91% of runs regardless. Without disclosure at least one other agent enters the death room in 95% of runs, and disclosure lowers this only to 81% because warnings are often challenged. Telling all agents that one of them knows the death room collapses deception in six of eight models; telling the informed agent that the room is permanently fatal moves only one. Applying an honesty steering vector to an open-weight model nearly eliminates deception without impairing play. The environment and game provide a testbed for measuring deceptive behavior and comparing interventions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.