PopGen: Generating Semantically Diverse Populations with LLMs for Zero-shot coordination
Abstract
In multi-agent reinforcement learning (MARL), exposing a cooperator to partners with diverse strategies during training is key to better collaboration at test time. However, existing methods typically rely on trajectory-level proxies to measure diversity, which may label two distinct trajectories of the same strategy as diverse. Our key insight is instead to leverage LLMs' ability to describe behaviorally distinct agents in environments, similar to how humans can recognize high-level strategies. To this end, we introduce , a framework that synthesizes semantically diverse agents from scratch using LLMs. first generates a diverse set of textual descriptions of agent behaviors, then translates each behavior into a code policy that can be executed within the environment. We show that training a cooperator against a population achieves robust zero-shot cooperation performance across four collaborative environments, and that the resulting policies adapt to diverse partners. Additionally, we measure how diversity scales with larger populations and provide insights on how population diversity impacts cooperation and how it can conflict with the quality of the agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.