SHRED: Self-Play for Heterogeneous Road-Users without Expert Demonstrations
Abstract
Realistic multi-agent traffic simulation requires modeling the diverse behaviors and dynamics of vehicles, pedestrians, and cyclists interacting in shared environments. Current approaches predominantly rely on imitation learning from large-scale human driving datasets, achieving strong realism at the cost of demonstration-data dependence and computational overhead. We present SHRED, a framework for training heterogeneous populations of vehicles, pedestrians, and cyclists entirely through multi-agent reinforcement learning from random initialization, without trajectory supervision, logged-trajectory replay, or a pretrained imitation-policy anchor. Our approach combines type-specific dynamics, rewards, and procedural scenario generation, including randomized spawning, map-conditioned goal generation, and per-agent randomization of kinematic and reward parameters, enabling decentralized recurrent policies to produce diverse and coordinated behaviors. All agent types are trained jointly within a unified heterogeneous self-play pipeline despite their distinct dynamics and action spaces. Evaluated on the Waymo Open Sim Agents Challenge (WOSAC), SHRED jointly controls vehicles, pedestrians, and cyclists in closed loop. Component ablations favor type-specific objectives as well as dynamics, with uneven benefits across road-user types. SHRED achieves, to our knowledge, state-of-the-art performance on the official InterPlan benchmark, which tests interaction-heavy closed-loop planning. A policy trained on nuPlan map geometry retains comparable realism when evaluated on Waymo, supporting transfer across map sources. Our results suggest that heterogeneous procedural self-play can serve as a scalable and complementary paradigm to data-driven traffic simulation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.