Towards Learning Representations of Policies in Two-Player Zero-Sum Imperfect-Information Games
Abstract
Search-based methods for two-player zero-sum imperfect-information games reason over public belief states, whose evolution depends on the players' policies. Scaling these methods therefore calls for compact representations of policies, not just of states. We study self-supervised learning of policy embeddings: low-dimensional vectors that summarize a policy's strategic behavior and, ideally, can be decoded back into a playable policy. We make three contributions. First, we introduce three ways to generate diverse policy populations. Second, we compare six representation methods: weight-space autoencoders, behavioral encoders trained with contrastive or imitation objectives, and the conditioning embeddings of a NeuPL population network. Third, we propose a suite of downstream tasks. They range from linear probes of payoff and exploitability to zero-shot best response and equilibrium search carried out directly in embedding space. On Kuhn and Leduc poker, behavioral encoders cut exploitability-prediction error by up to 87% while weight-based ones fail, and policy optimization in latent space improves exploitability on Leduc. Our results point to behavior-grounded, decodable embeddings as a building block for latent-space planning in imperfect-information games.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.