Coding Agents are Zero-shot Coordinators
Abstract
We focus on zero-shot coordination (ZSC), where independently developed agents must cooperate with unseen partners at test time (e.g., autonomous driving). Self-play (SP), a paradigm in which agents learn by interacting with copies of themselves, leads them to adopt specialised conventions that fail to generalise to new partners. Existing ZSC methods rely on restrictive assumptions and perform poorly on benchmarks such as Overcooked V2 and Yokai, achieving poor inter-seed cross-play (IXP), which measures mutual compatibility across independent runs of the same algorithm. Recent demonstrations of large-scale LLM coordination, together with rapid advancements in their coding capabilities, motivate a different approach: using general-purpose LLMs to directly generate ZSC policies as code. We find that coding agents, LLMs equipped with an off-the-shelf coding harness and runtime environment for iterative policy refinement, achieve state-of-the-art IXP on both Overcooked V2 and Yokai. Their policies also remain compatible across model families and novel Overcooked V2 configurations. Beyond ZSC, we consider a setting in which agents may pre-agree on conventions before deployment but are prohibited from joint training, as in self-play. Subsequently, we find that coding agents generate policies that reliably follow the H-group set of conventions used by human players in Hanabi. Taken together, our results show that coding agents significantly improve over existing ZSC methods, offering a promising route toward robust, adaptable, and human-compatible coordination.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.