acceptodds
Under review as a conference paper at ICLR 2027

SharedIntentBench: A Benchmark for Agents That Coordinate for Independent Principals

Abstract

We introduce SharedIntentBench, a benchmark for AI agents that coordinate deals for different people, such as booking a meeting or making a purchase. Each person gives their own agent private requirements and can change them mid-negotiation. The benchmark has 24 two-party task families in four domains. In eight of them, one person changes their requirements while the agents work. An automatic checker sees everyone's private requirements and approvals. The checker scores two opposite failures. First, the agents can close a deal on an offer made before the change, although nothing shows that the person still agrees. We call this a stale completion. Second, the agents can end with no deal although an option acceptable to everyone remains. We call this a walkaway. A deal is correct only if everyone accepts its terms and every approval is current. A no-deal is correct only if no acceptable option remained. Checking only whether a deal was completed misses the first failure. Counting every no-deal as safe misses the second. Across eight models and 3,600 trials, the share of correct endings ranges from 33.5% to 97.5%. Walkaways are about three times as common as stale completions, and they account for most failures of GPT-5.6 Sol and Qwen3-235B. Scoring the same trials with only one of the two checks reorders the models. In controlled tests with open-weight models, an update that changes no requirement still induces agents to close deals on old offers. Blocking such deals cuts stale completions from 40 to 3 in 768 trials, but walkaways remain the most common failure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.