Lookahead-Guided Closed-Loop Fine-Tuning for Multi-Agent Traffic Simulation
Abstract
Traffic simulation is crucial for the development of autonomous driving systems. Recent autoregressive approaches formulate this task as sequential motion token prediction, yet inherently suffer from covariate shift between open-loop training and closed-loop deployment. Closed-loop fine-tuning (CLFT) mitigates this issue by sampling pseudo expert trajectories. However, most existing token-level methods rely on greedy single-step sampling, and locally closest matches do not guarantee global trajectory alignment. Inspired by multi-token prediction in large language models, we propose Lookahead-Guided Closed-Loop Fine-Tuning (LG-CLFT), a simple yet effective framework that enables long-horizon matching for pseudo expert trajectory construction. Specifically, we introduce auxiliary multi-step prediction heads to approximate future motion token distributions during open-loop pretraining. During CLFT, we leverage this long-horizon prediction capability to perform beam search over future motion token sequences, evaluating their overall alignment with expert trajectories. This lookahead mechanism enables sampling superior pseudo expert trajectories within the model policy distribution. On the WOSAC 2025 leaderboard, LG-CLFT achieves a Realism Meta score of 0.7854 and a minADE of 1.2865, outperforming prior supervised CLFT methods in both simulation realism and trajectory alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.