Hierarchical Safety-Oriented Dense RL with Refined Credit Assignment for Tokenized Multi-Agent Traffic Simulation
Abstract
High-fidelity simulation of safety-critical traffic interactions is essential for developing and evaluating autonomous driving systems. Tokenized autoregressive models achieve strong average realism in multi-agent traffic simulation through large-scale supervised pretraining, but safety-critical events such as collisions are rare in natural driving data, providing limited supervision for learning these interactions. Uniformly sampled reinforcement learning (RL) post-training also allocates many updates to ordinary scenarios, agents, and timesteps that provide little safety information, limiting the efficiency of safety-oriented learning. We propose a hierarchical safety-oriented dense learning framework for RL post-training of tokenized multi-agent traffic simulation models. Episode-level densification prioritizes scenarios whose current-policy probe produces a collision, while state-level masks localize updates to vehicle agent-time positions with a collision or physical near miss in at least one same-scenario rollout, retaining aligned safe outcomes for comparison. Our refined credit assignment retains physical rewards at agent-step granularity and compares aligned agent-time positions across same-scenario rollouts, yielding critic-free group-relative advantages with explicit temporal-agent resolution. Experiments with SMART on the Waymo Open Motion Dataset show that, across three training seeds, our method reduces mean collision rate from 3.394% to 2.914% relative to the uniformly sampled RL baseline using the same refined credit assignment, while maintaining a comparable average displacement error, achieving a favorable balance between safety-oriented alignment and per-rollout trajectory fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.