APEX: Robust Driving Self-Play through Adaptive Population Exploration
Abstract
Large-scale driving self-play can learn robust policies through closed-loop interaction without human driving demonstrations. Yet existing methods use shared policies and fixed traffic-behavior distributions, causing driving policies to overfit to training-time behavioral regularities and generalize poorly to unseen traffic behaviors. We introduce Adaptive Population Exploration for Self-Play (APEX), which reframes driving self-play as a population-level multi-objective learning problem. Specifically, APEX separates the target driving policy, termed the focal policy, from a behavior-conditioned traffic policy, while making the distribution of traffic behaviors learnable. While the two policies interact and optimize their respective objectives, a population constructor uses focal-policy feedback to adapt the distribution of traffic conditions, thereby changing which self-interested behaviors the focal policy encounters without directly controlling vehicle actions. To focus this adaptation on policy weaknesses rather than intrinsic scene difficulty, we use the focal critic's initial value estimate as a baseline for realized focal performance. We show that this centering preserves the expected optimization direction of the population objective under our training protocol. By alternating policy learning and population adaptation, APEX tracks evolving weaknesses. Experiments show improved robustness and generalization to unseen traffic behaviors while maintaining competitive performance under standard self-play.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.