acceptodds
Under review as a conference paper at ICLR 2027

Coding Agents are Zero-shot Coordinators

Abstract

We focus on zero-shot coordination (ZSC), where independently developed agents must cooperate with unseen partners at test time (e.g., autonomous driving). Self-play (SP), a paradigm in which agents learn by interacting with copies of themselves, leads them to adopt specialised conventions that fail to generalise to new partners. Existing ZSC methods rely on restrictive assumptions and perform poorly on benchmarks such as Overcooked V2 and Yokai, achieving poor inter-seed cross-play (IXP), which measures mutual compatibility across independent runs of the same algorithm. Recent demonstrations of large-scale LLM coordination, together with rapid advancements in their coding capabilities, motivate a different approach: using general-purpose LLMs to directly generate ZSC policies as code. We find that coding agents, LLMs equipped with an off-the-shelf coding harness and runtime environment for iterative policy refinement, achieve state-of-the-art IXP on both Overcooked V2 and Yokai. Their policies also remain compatible across model families and novel Overcooked V2 configurations. Beyond ZSC, we consider a setting in which agents may pre-agree on conventions before deployment but are prohibited from joint training, as in self-play. Subsequently, we find that coding agents generate policies that reliably follow the H-group set of conventions used by human players in Hanabi. Taken together, our results show that coding agents significantly improve over existing ZSC methods, offering a promising route toward robust, adaptable, and human-compatible coordination.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.