How Much Do LLM Negotiations Reveal Beyond What Their Agreements Require?
Abstract
LLM agents negotiate for users and act on the users' private preferences. The agents' offers and agreements let the other side infer those preferences. Some disclosure is the price of a good agreement. We ask how much an LLM negotiation reveals beyond that price, with the price measured by welfare, the quality of the agreement. Existing negotiation and agent-privacy benchmarks cannot answer this question, because what a negotiation must reveal depends on the agreement the negotiation reaches and on what the other side already knows. We define the delegation frontier: the least information that reaching a given welfare must reveal about each user to the others, even with an ideal mediator. A decoder learns to guess preferences from the negotiation record, and the decoder's score is a lower bound on disclosure. We subtract an upper bound on the frontier and obtain an excess estimate. Decoder error can only lower this estimate, so a positive estimate is conservative. An exact census of a reduced bargaining game over every preference profile shows what the decoder misses. We audit seven models in four-party bargaining and four models in bilateral procurement. Estimated excess is positive in all 33 original-prompt conditions. In the census, at least 2.5 of at most 8 disclosed bits are excess, and 79–99% of the excess lies in the agreement rather than in the offers exchanged. The excess is not inevitable: a scripted bargainer usually reaches the same welfare with less disclosure. Telling agents to accept reasonable offers lowers disclosure without detectable welfare loss. However, the instruction transfers the decision to the first proposer, and the first proposer gains surplus in two of three models validated on fresh profiles. Decoder error also varies across policies. As a result, decoded bounds overstate seven of eight intervention effects in the census and reverse two of four policy rankings. Choosing among interventions therefore requires checking the auditor.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.