World2Policy: Continual Policy Improvement via Scenario Curriculum in Interactive Gaussian Worlds
Abstract
Interactive Gaussian worlds enable photorealistic closed-loop simulation for autonomous driving. However, sustained policy improvement requires addressing two coupled challenges: how to resolve heterogeneous safety constraints during policy learning, and which scenarios should support continual policy update. A monolithic policy often entangles avoidance behaviors across dynamic actors, static obstacles, and road boundaries, while fixed scenario distributions yield diminishing returns as the policy improves. We present **World2Policy**, a co-evolutionary framework that unifies policy optimization with failure-conditioned scenario generation. In the *Policy Optimization Engine*, three constraint-specialized experts decouple safety modeling across foreground dynamics, static hazards, and drivable-area constraints through constrained reinforcement learning, while a privileged selector evaluates their complementary proposals to balance safety, progress, and comfort. In the *Environment Engine*, closed-loop rollouts expose residual failures, which guide the synthesis of physically plausible scenarios near the policy's capability boundary and update the training curriculum for policy optimization. Experiments on the closed-loop HUGSIM benchmark demonstrate that integrating World2Policy with UniAD, VAD, and LTF achieves HD-Score gains of +4.9, +6.1, and +6.9 points, respectively. Crucially, the co-evolutionary loop continually advances the policy's capability and sustains improvements across successive training rounds.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.