Learning to Coordinate Through Strategic Game Slices
Abstract
Coordination under private information requires agents to infer partners’ preferences and select actions using those inferences. We study these capabilities in a finite negotiation game with private preferences and mixed incentives. A reference-policy oracle supplies posterior judgments and counterfactual action values, enabling matched diagnostic interventions. For Qwen3-4B-Instruct-2507, these probes reveal errors in both preference inference and action selection: supplying an oracle judgment does not reliably eliminate action errors. We use the same oracle to construct game slices and train a shared model on direct action selection, preference inference, and planning. With approximately 6M generated response tokens, joint slice training retains 47 of 72 meetings in a 24-game, four-agent sequential calendar suite adapted from CalBench, compared with 37 for both the unchanged model and complete-game self-play. These results provide initial evidence that targeted supervision from controlled strategic games can improve coordination in an environment with different interaction mechanics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.