acceptodds
Under review as a conference paper at ICLR 2027

Learning to Bargain through an Allocation-Sensitive Value Bridge

Abstract

Learning to bargain within a larger sequential decision problem requires agents to identify which exchanges to make and which proposals others will accept. When the value of an exchange is endogenous to the surrounding task, agents face sparse and high variance reward signals during bargaining. We introduce a hierarchical reinforcement-learning framework that addresses this by learning an allocation-sensitive value bridge that provides immediate post-bargaining feedback without prescribing agent behavior or relying on negotiation demonstrations. We test this framework in a four-player variant of the board game Settlers of Catan with two independently seeded self-play league lineages. Our experiments demonstrate that agents trained with the value bridge win more often against common opponents, reach agreement more often, and rely more on negotiated player exchange. This performance improvement is attributable to gains in bargainer ability. Ablations show that without the value bridge the bargainer loses selectivity, performs no better than a random bargainer, and the complete agent wins well below par. These results show that economically recognizable bargaining behavior can emerge from task-oriented self-play using a learned interface between negotiation and long-horizon control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.