Voluntary Collusion with Secret Tools in Competing LLM Agents
Abstract
Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents can voluntarily accept it for strategic advantage. We introduce an empirical framework that separates the voluntary adoption of a collusive capability from its subsequent use in two multi-agent environments: Liar's Bar, a competitive deception scenario, and Cleanup, a mixed-motive resource-management environment. Across twelve models and six offer framings, ten models accept both secret tools at rates of at least 98% under an explicitly unfair offer, often while acknowledging the unfairness in their explanations. The only two models that refuse an explicitly unfair offer accept the same capability when it is described neutrally, as do four additional frontier models. When the tools operate as offered, alliance members coordinate against other agents, producing persistent score gaps and lower reward equality under the tested tool conditions. Our work establishes voluntary collusion adoption as a distinct evaluation target, shifting attention from whether agents can collude to whether they choose to acquire collusive capabilities, and shows that refusing an explicitly unfair offer does not imply refusal of the capability itself.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.