Scalable Rollout Generation for Generative Driving Simulators
Abstract
Generative simulators are emerging as a promising foundation for post-training of autonomous driving policies. However, scaling training in these environments is limited by the cost of sensor simulation and the sequential dependency between observation rendering and policy inference. The simulator waits for the policy to predict the next ego pose, while the policy waits for the resulting observation before replanning. This ordering introduces idle time and limits rollout throughput. Inspired by speculative decoding, we introduce Speculative Rendering, a rollout-generation framework that speculates future simulation states ahead of planner execution. Our approach renders future observations in parallel with the planner, and employs a verifier to determine whether the speculated state remains sufficiently close to the realized state. When the divergence exceeds an acceptance threshold, the speculated renders are discarded and the simulator state is resynchronized with the target planner. We evaluate Speculative Rendering for GRPO post-training of AutoVLA using WorldEngine, a generative driving simulator. Speculative Rendering reduces mean simulation-step latency from to ms and lowers rollout-generation cost from to H100 GPU-hours, a reduction. The resulting policy trained on speculative rollouts improves PDMS on the challenging navtest-failure benchmark by and remain within of the performance obtained using standard sequential rendering.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.