Finding and Closing Coverage Holes in Vision-Language Motion Planners
Abstract
Vision-language motion planners have advanced rapidly, yet single-trajectory accuracy can conceal missing feasible alternatives in imitation-trained planners. We define a coverage hole as a state–direction pair with verified trajectory witnesses, but no feasible sample from a reference planner at a fixed budget. Our verifier checks drivable-area compliance and collision-freeness against recorded agents, not traffic-rule compliance. We diagnose candidate holes using retimed cross-scene trajectory prototypes. The same witnesses supply augmentation targets, while a separate human-supervised detector selects the evaluation set. Neither training nor detection requires HD-map or maneuver annotations. Witness-target augmentation is combined with rehearsal and prefix-scheduled decoding (PSD), which allocates inference samples across direction prefixes and free generation using a scene-independent schedule. On nuScenes, we evaluate a 10B vision-language planner on 254 reference-defined hole pairs at 200 training-split states (about half of the pairs received a training target) and on validation and held-out scenes. Augmentation raises closure@128 from 46.5% to 85.0%, while minADE@8 decreases from 2.39 to 1.72 m. Adding PSD reaches 90.9% closure [86.0, 95.1] at 1.91 m minADE and achieves 70% closure with about 29 samples. The off-the-shelf planner remains below 70% at 128. Validation closure under the official HD map reaches 96.0% versus 38.1%, with gains persisting under a lane-direction check. GT-only fine-tuning is more accurate (1.26 m) but recovers fewer reference-defined holes, and a matched-budget control with random witness targets closes most holes as well, so hole targeting mainly buys sampling efficiency at inference. These results show improved finite-budget candidate coverage under an open-loop verifier, with an accuracy trade-off relative to GT-only training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.