acceptodds
Under review as a conference paper at ICLR 2027

Vault: One-Step Latent Generation with Positive-Anchored Rewards for Autonomous Driving

Abstract

End-to-end autonomous driving must balance multi-modal maneuver generation against real-time inference constraints. Diffusion planners capture diverse behaviors but their iterative denoising is too slow for safety-critical deployment, while one-step alternatives still imitate only the single expert demonstration, leaving neither diversity nor quality certified. To address this, we propose Vault, a framework that couples one-step latent generation with sample-based reinforcement learning guidance. Vault generates candidates in a single forward pass by drifting in a VAE latent space, built on a V-JEPA 2.1 visual representation pretrained on the full NAVSIM training imagery. During training, a reward-gated positive pool collects the full-score trajectories that the model itself discovers under the official evaluation and anchors the drift targets to these certified samples. This closes a reinforcement-learning loop purely through sample-based target construction, without policy gradients, learned reward models, or extra interaction: the generator imitates the frontier of its own certified successes, so quality and multi-modality improve together under a single fixed sampling configuration. A learned scorer predicting the official score and its sub-metrics selects the executed trajectory, and all training-only modules are discarded at inference. On the NAVSIM benchmarks, Vault achieves state-of-the-art closed-loop performance in real time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.