Z2SP: Turning Zero-Shot Coordination into Self-Play via Partner-Specific Test-Time Adaptation
Abstract
Zero-shot coordination (ZSC) requires independently trained agents to coordinate with unseen partners in cooperative multi-agent reinforcement learning (MARL). Most existing methods learn policies that generalize broadly before deployment. Once a particular partner appears, however, the optimization focus of ZSC shifts from distribution-level generalization to online partner-specific policy estimation and adaptation. When the partner's self-play return exceeds that of the initial ZSC pairing, the current partner policy provides a concrete adaptation target: moving the ego policy toward it can steer coordination toward a self-play-like mode, narrowing the ZSC-to-SP gap and improving ZSC performance. We present Z2SP, which constructs a finite-sample approximation of the inaccessible partner policy over encountered observations from online partner-view observation–action samples and directly adapts the ego actor toward this target. This unifies partner-policy estimation and ego-policy adaptation without a separate partner model. All adaptation uses samples collected only after the initial zero-shot encounter. For persistent adaptation, Z2SP retains its adaptation state across episodes, reuses accumulated samples through delayed history replay, and employs closed-loop meta-initialization to mitigate early adaptation cost. Across five OvercookedV2 layouts with unseen partners from four MARL families, Z2SP recovers 91.8% of the initial ZSC-to-SP gap after 40 episodes of persistent adaptation and averages 98.0% recovery over the final 20 episodes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.