acceptodds
Under review as a conference paper at ICLR 2027

Cross-Environment Cooperation with Decoupled Policy and Value Learning

Abstract

Zero-shot coordination (ZSC) enables ego agents to collaborate with unseen partners and is essential for human–AI coordination with diverse human behaviors. Recent work has improved coordination generalization by training agents across diverse environments. However, cross-environment training can still be limited by overfitting. Prior work in single-agent reinforcement learning shows that shared policy-value representations can hinder generalization by introducing environment-specific information into the policy representation. We hypothesize that shared policy-value representations also limit generalization in cross-environment ZSC. To investigate this hypothesis, we incorporate IDAAC, which decouples policy and value learning, into Cross-Environment Cooperation and introduce DCEC. Our analysis shows that DCEC selects actions more consistently across similar states from different environments, while encoding less environment-specific information in its policy representation. Moreover, DCEC exhibits higher value-gradient GSNR than CEC at larger training batch sizes, indicating more consistent value-learning gradients under larger-batch training. Across Dual Destination and Overcooked-AI, DCEC improves coordination with unseen partners and generalization to unseen environments, while also achieving strong coordination performance with real humans.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.