acceptodds
Under review as a conference paper at ICLR 2027

Towards Learning Representations of Policies in Two-Player Zero-Sum Imperfect-Information Games

Abstract

Search-based methods for two-player zero-sum imperfect-information games reason over public belief states, whose evolution depends on the players' policies. Scaling these methods therefore calls for compact representations of policies, not just of states. We study self-supervised learning of policy embeddings: low-dimensional vectors that summarize a policy's strategic behavior and, ideally, can be decoded back into a playable policy. We make three contributions. First, we introduce three ways to generate diverse policy populations. Second, we compare six representation methods: weight-space autoencoders, behavioral encoders trained with contrastive or imitation objectives, and the conditioning embeddings of a NeuPL population network. Third, we propose a suite of downstream tasks. They range from linear probes of payoff and exploitability to zero-shot best response and equilibrium search carried out directly in embedding space. On Kuhn and Leduc poker, behavioral encoders cut exploitability-prediction error by up to 87% while weight-based ones fail, and policy optimization in latent space improves exploitability on Leduc. Our results point to behavior-grounded, decodable embeddings as a building block for latent-space planning in imperfect-information games.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.