Wider Coverage, Better Plan: Admissible Alternatives for End-to-End VLA Planning
Abstract
In autonomous driving, most scenes admit several routes to proceed, and a planner, often a Vision-Language-Action (VLA) model today, must recognise these options yet commit to one of them. Recent works enable VLA planners to sample extensive trajectories with generative heads, and on challenging benchmarks, the strongest methods pick one of thousands of candidates with a learned scorer. Can the VLA planner generate one or a small set of trajectories, instead of thousands of alternatives, that already contain high-quality options? We argue that supervised fine-tuning (SFT) on an offline dataset, where a single drive trajectory is sampled per scene, collapses the planner onto one mode. As such, we propose to train the planning head of a VLA planner on admissible alternative plans (AAP), additional plans per scene that a rule-based scorer accepts. Applied to Qwen-Drive-1.0 on the NAVSIM v2 navhard benchmark, AAP improves the single plan from 0.315 to 0.373 EPDMS, a gain that holds when both planners drive at equal speed, and turns near-identical samples into distinct options whose best of 16 rises from 0.359 to 0.614. A second stage of reinforcement learning makes the samples more reliable, and a speed-matched control shows that much of its gain comes from more cautious driving. We advocate evaluating planners by the diversity and coverage of their samples, not only by a single benchmark score. Code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.