acceptodds
Under review as a conference paper at ICLR 2027

Entity-Wise Factorization Enables Generalist Agents in Hanabi

Abstract

Self-play lets multi-agent reinforcement learning agents cooperate well with copies of themselves, but the arbitrary conventions they invent along the way rarely transfer to a new partner, human or otherwise. This is exemplified in Hanabi, a cooperative card game for 2 to 5 players widely used to probe theory of mind. Until now, -player game has been considered separately, with multiple works focusing only on 2- or 3-player games. In this work, we propose a generalist Hanabi agent that can play and be trained on any -player game jointly. To enable this, we require a common state representation agnostic to the number of players. Inspired by UPDeT (Hu et al., 2021b), a successful example of factorizing an observation into per-entity tokens so that one set of parameters can handle a variable number of entities, we propose FHA (factorized Hanabi agent), a simplified and fast factorization architecture. We then perform a comprehensive analysis of self-play, other-play and off-belief learning trained with our proposed architecture, jointly and independently across players. In doing so, we observe that factorization and joint training can improve cross-play performance of self-play agents, because the agent learns a general policy that it applies regardless of the number of players. Thus, by avoiding the overspecialization issues encountered by self-play agents, self-play can be a viable training method for cross-play agents. Finally, the learned policy is transferable, as FHA is able to train on a subset of settings and zero-shot transfer at deployment to unseen settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.