Learning End-to-End Driving Policies Without Expert Demonstrations
Abstract
End-to-end autonomous driving has achieved state-of-the-art performance on benchmarks and real-world deployments. Its standard training recipe, however, is expensive across all existing paradigms: imitation learning requires fleet-scale driving logs whose behavior distribution only thinly covers the safety-critical tail; closed-loop reinforcement learning in a photorealistic simulator, while offering diverse behaviors, is bottlenecked by sample-efficiency, per-step cost of rendering, and a forward pass through a large vision backbone. Self-play in vectorized simulators changes the economics: at millions of rollout steps per second, one can obtain a state distribution with natural diversity in collisions, near-misses, and recoveries that no driving log contains, and a highly reactive policy emerges from a basic driving-quality reward. Our approach exploits this asymmetry by decoupling learning to drive from learning to see. We pretrain a single policy by self-play, then align a pretrained vision backbone to its latent space through the action Kullback-Leibler divergence and a batch-relational low-rank structural loss. The action target comes from the self-play policy, so alignment never supervises against a logged trajectory: a paired dataset of (image, scene-state) frames suffices, with no need for the curated expert demonstrations that imitation pretraining is built on. The proposed framework establishes a novel learning paradigm that substantially reduces the cost of end-to-end driving policy learning, while the policy learned significantly outperforms every published imitation-based end-to-end method on photorealistic 3D Gaussian splatting closed-loop benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.