FPD-Drive: Future Policy Distillation for End-to-End Autonomous Driving
Abstract
End-to-end autonomous driving models are typically trained on real-world driving logs, which record only the trajectory actually executed in each scene, providing limited supervision for alternative valid behaviors. Multi-target supervision can mitigate this limitation, but broader trajectory coverage makes safe candidate selection more challenging. We propose FPD-Drive, an end-to-end planner that distills privileged future knowledge into a visual driving policy to enrich trajectory supervision and improve safety-aware candidate selection. Specifically, Future-Privileged Multi-Target Distillation transfers diverse, high-quality trajectory targets from a teacher with access to recorded future observations, extending student supervision beyond the single recorded human trajectory. Safety-Aware Score Correction further corrects ranking errors between nearby safe and unsafe candidates and raises underestimated safety scores of safe candidates. At deployment, the student requires neither future observations nor teacher inference. Our base model achieves 94.9 PDMS on NAVSIM v1 navtest and 55.9 EPDMS on NAVSIM v2 navhard. The scaled FPD-Drive-Pro improves these scores to 95.5 PDMS and 58.1 EPDMS, respectively, achieving state-of-the-art performance among the compared methods on both benchmarks. Zero-shot closed-loop evaluation on HUGSIM further yields an HD-Score of 45.1 and RC of 55.3.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.