Cross-Environment Cooperation with Decoupled Policy and Value Learning
Abstract
Zero-shot coordination (ZSC) enables ego agents to collaborate with unseen partners and is essential for human–AI coordination with diverse human behaviors. Recent work has improved coordination generalization by training agents across diverse environments. However, cross-environment training can still be limited by overfitting. Prior work in single-agent reinforcement learning shows that shared policy-value representations can hinder generalization by introducing environment-specific information into the policy representation. We hypothesize that shared policy-value representations also limit generalization in cross-environment ZSC. To investigate this hypothesis, we incorporate IDAAC, which decouples policy and value learning, into Cross-Environment Cooperation and introduce DCEC. Our analysis shows that DCEC selects actions more consistently across similar states from different environments, while encoding less environment-specific information in its policy representation. Moreover, DCEC exhibits higher value-gradient GSNR than CEC at larger training batch sizes, indicating more consistent value-learning gradients under larger-batch training. Across Dual Destination and Overcooked-AI, DCEC improves coordination with unseen partners and generalization to unseen environments, while also achieving strong coordination performance with real humans.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.