acceptodds
Under review as a conference paper at ICLR 2027

Enhanced Simulation Environment Based on Distribution Matching and Large Language Model-Guided Reinforcement Learning

Abstract

Autonomous driving simulation environments often suffer from systematic statistical drift during long-term closed-loop operation, where the behavior of background vehicles gradually deviates from real-world traffic statistics. Existing enhancement methods either rely on rule-based models with limited statistical fidelity or retain online large language model (LLM) calibration during simulation, which introduces prohibitive computational overhead for real-time multi-vehicle control. To alleviate the above limitations, this paper proposes Enhanced Simulation environment based on Distribution matching and Large language model Guided reinforcement learning (ESDLG), a two-stage simulation enhancement framework for background-vehicle policy learning. Rather than introducing a new divergence measure or a new reinforcement learning algorithm, ESDLG introduces χ²-divergence-based state occupancy matching and uses LLM-guided policy learning for long-term traffic simulation enhancement. In the first stage, inspired by state occupancy matching formulations, ESDLG introduces a χ²-divergence dual optimization module into PPO; PPO learning is designed to simulate and control vehicles after training, to provide a real-data-driven distribution-matching prior without requiring expert action labels. In the second stage, real trajectories are clustered into three distinct driving styles, and three style-specific PPO policies are fine-tuned using LLM-guided exploration. The LLM guidance is used only during training and distilled into lightweight policy networks; during deployment, the LLM is completely removed from the simulation loop. Experiments conducted in highway-env with HighD data show that ESDLG achieves balanced traffic-statistic fidelity, obtaining the best spacing MAPE and the second-best Hellinger distance for speed-distribution matching among the compared methods. ESDLG also preserves competitive coarse-grained driving style diversity while reducing the per-step inference time to 15.9 ms, corresponding to a 5,424× speedup over online LLM calibration. Additional evaluations on traffic density and style ratio further indicate that ESDLG remains robust under varying traffic densities and supports soft scenario-level controllability through policy assignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.