acceptodds
Under review as a conference paper at ICLR 2027

Do Agents Keep Their Word? Promise-Breaking in Multi-Agent LLMs

Abstract

Monitoring what large language model (LLM) agents say they will do is an active field of alignment research for multi-agent systems. However, we know little about whether the agents keep their promise. We evaluate ten frontier models on six social-dilemma games with two protocols. In the exogenous protocol the agent is assigned a public commitment under seven conditions, and we compare its deviation rate with the rate it would reach by ignoring the commitment. In the endogenous protocol the agent privately plans, publicly announces its own commitment, and then acts. We find that promise-breaking depends on the model and the game together, with their interaction accounting for 48% of the variance across model–game cells, twice as much as the model alone. A single instruction to honor public statements reduces deviation from an assigned statement to under 2% on binary games, but self-made promises are still broken eight times as often. Declaring announcements binding has a similar effect, while declaring them cheap talk changes little, which suggests that models treat announcements as non-binding unless told otherwise. In repeated play, promise-breaking falls after the first round and individual models then diverge: over twenty-five rounds, one stops breaking entirely and another breaks more each round. Furthermore, when an agent breaks a promise, its reasoning usually does not mention the announcement. We score this with four LLM judges validated against three human annotators and find no clear relation between a model's awareness and how often it breaks its promises. These results suggest that multi-agent systems should state explicitly that announcements are to be honored, and should still expect some promises to be broken.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.