Adaptive Multi-Objective Metric-Oriented Policy Optimization for Multi-Agent Traffic Simulation
Abstract
High-fidelity multi-agent traffic simulation is an indispensable foundation for benchmarking, validating, and training autonomous driving systems. Standard imitation learning (IL) and supervised fine-tuning (SFT) suffer from two foundational bottlenecks: covariate shift in closed-loop autoregressive rollouts and metric misalignment, where maximum-likelihood token objectives diverge from non-differentiable downstream requirements such as kinematic safety, collision avoidance, road adherence, and interactive realism. Recent reinforcement fine-tuning (RFT) methods like SMART-R1 introduced Metric-oriented Policy Optimization (MPO) using empirical advantage formulations . However, existing MPO paradigms are constrained by three severe pathologies: (i) reliance on a static scalar threshold that induces negative advantage collapse in challenging interactive scenes while over-rewarding trivial scenarios, (ii) monolithic scalar rewards that collapse under multi-objective Pareto conflicts (sacrificing diversity and traffic throughput for over-conservative creeping), and (iii) prohibitive closed-loop evaluator latency scaling as . To overcome these fundamental limitations, we propose Adaptive Multi-Objective Metric-Oriented Policy Optimization (AM-MPO), a principled, scalable, and safety-aware post-training framework for multi-agent closed-loop traffic simulation. AM-MPO introduces four synergistic contributions: (1) Adaptive Threshold Advantage Estimation, which dynamically combines exponential moving averages with context-conditioned percentile calibration over scene complexity descriptors (agent density, topological curvature, interaction graph intensity) to eliminate negative advantage collapse; (2) Curriculum-Driven Multi-Objective Reward Decomposition, decoupling realism, safety, diversity, efficiency, and kinematic comfort through dynamic weight scheduling; (3) Safety-Aware Constrained Policy Optimization, enforcing strict kinematic and legal safety boundaries via primal-dual Lagrangian relaxation to prevent reward hacking; and (4) Surrogate Reward Distillation & Dense Temporal Credit Assignment, utilizing a lightweight graph surrogate model to slash evaluator computational overhead by while providing dense step-level temporal feedback. Evaluated on the large-scale Waymo Open Motion Dataset (WOMD) validation benchmark, nuPlan, and Argoverse 2 across model scales from 7M to 1B parameters, AM-MPO establishes a new state of the art: improving Realism Meta to 0.7761, increasing Collision-Free Rate to 97.31%, expanding Trajectory Diversity by +5.8%, and unlocking controllable, diverse rollout modes (conservative, balanced, aggressive-yet-safe) without manual heuristic retuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.