acceptodds
Under review as a conference paper at ICLR 2027

TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate

Abstract

Negotiation is a central mechanism of economic exchange, shaping outcomes in markets, procurement, labor agreements, and resource allocation. It also provides a canonical testbed for agentic language models, requiring multi-turn interaction under hidden preferences, strategic communication, and binding constraints. These properties make negotiation difficult to evaluate: unlike math or code, it has no intrinsic verifier. Existing LLM negotiation evaluations rely on LLM versus LLM interaction or aggregate outcome metrics such as deal rate, leaving failures opaque. We introduce **TERMS-Bench** (**T**estbed for **E**conomic **R**easoning in **M**ulti-turn **S**trategy), a Bayesian-game framework that makes the environment itself the verifier by specifying the counterpart's latent type, policy, and payoff structure. We instantiate this framework in bilateral price negotiation, where the counterpart's private state and simulator policy are hidden from the agent but observable to the evaluator. This turns the counterpart from a black-box opponent into a diagnostic instrument, enabling agent-attributable failure analysis and oracle reference policy defined optimality gaps. Evaluating 13 LLM agents spanning frontier systems from major providers, **TERMS-Bench** turns negotiation evaluation from aggregate ranking into actionable diagnosis: *where* agents fail, *why* they fail, and *what* to strengthen. Empirically, frontier models saturate deal rate yet diverge in surplus, cue use ability, belief calibration, and compliance, revealing agent-specific bargaining bottlenecks masked by prior benchmarks. Stress tests of the counterpart preserve the core findings, and a model-versus-model round-robin shows that the capability measured by the simulator largely carries over to head-to-head play between models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.