RePT: Re-targeted Post-Training for Multi-Agent Traffic Simulation
Abstract
Traffic agents for closed-loop evaluation are trained by imitating logged driving, and reinforcement learning post-training is meant to repair the interactions imitation gets wrong: the merges and unprotected turns where a driver could yield or push ahead, but the log shows only what one driver did. We show that current post-training does not fix these scenes, and trace the reason to where its reward is placed. Existing methods reward each rollout for resembling recorded driving and apply that reward across the whole training set. Most of the gradient therefore goes to scenes the model already drives well. On the difficult scenes every rollout fails, so rewards are nearly equal, the update is weak, and the little signal left simply reinforces the least bad failure. RePT (Re-targeted Post-Training) makes two changes. It post-trains only on the scenes the supervised model scores lowest, chosen once before training. And because these scenes collapse onto one failing behavior before any alternative can succeed, it adds a diversity bonus: rollouts are clustered by their joint trajectories, since what must stay distinct is whether an agent yields or pushes ahead, each is paid for standing outside the others' clusters, and none is paid once its realism falls below the group. Alternatives thus stay in play until the realism reward can tell which one clears the interaction. On WOSAC 2025, RePT sets the highest realism reported and is the only post-training we test, ours or published, that improves the hardest scenes at all, with no change to architecture or inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.