Emergent Tool Use in the Hide-and-Seek Environment via an Adaptive Auto-Curriculum
Abstract
We study the emergence of coordinated multi-agent strategies in the quadrant configuration of Openai's hide-and-seek environment. Two teams, hiders and seekers, are trained purely through self-play under a zero-sum reward with no additional rewards on specific behaviours. To drive learning we combine self-play against a pool of past opponents with an adaptive auto-curriculum, inspired by Deepmind's Adaptive Agent research. Over the course of training we observe a sequence of distinct emergent phases: pursuit and evasion, door blocking, seeker ramp use, and a hider ramp-defense strategy. We monitor the emerge of these phase transitions and, in a controlled ablation across three random seeds, find that the adaptive curriculum reaches all four phases faster than a non-adaptive task sampler and with lower seed-to-seed variance. Relative to the phase onsets originally reported for the full hide-and-seek environment, our compact single-scenario setup reaches the final phase roughly 1.3 times faster (about 3M episodes), while following the same ordering of emergent behaviours.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.