Scorable Negotiation Self-Play for Natural-Language Reasoning and Understanding
Abstract
Can learning to negotiate improve language models beyond negotiation itself? We answer this question affirmatively with Language Model Self-Play via Scorable Negotiation Game (LSSG), a post-training framework that transfers skills from multi-turn bargaining to natural-language reasoning and understanding without any downstream reasoning labels. The key idea is to pair open-ended language interaction with continuous outcome rewards computed directly from transaction prices, with no LLM judge. LSSG warm-starts a single buyer–seller model by generalization-aware behavioral cloning on real dialogues, then refines it by stability-aware self-play, which combines payoff-driven updates with supervised retention and semantic diversity and emotional stability regularization from frozen pretrained scorers. Across seven benchmarks spanning commonsense reasoning, textual inference, logical reasoning, sentiment classification, and knowledge-intensive understanding in English and Chinese, LSSG achieves the highest accuracy among all compared methods in every one of the 14 model–benchmark settings on Qwen3-8B and Qwen3-14B, with macro-average gains of 4.61 and 3.14 percentage points over the vanilla models and up to 8.61 points on a single benchmark. Accuracy rises with every self-play iteration, and LSSG further raises win rate and payoff in two interactive negotiation games, also under adversarial social personas, establishing scorable negotiation as a judge-free training environment whose benefits extend well beyond bargaining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.