acceptodds
Under review as a conference paper at ICLR 2027

RethinkDrive: Supervision Targets for Lightweight Trajectory Refinement in End-to-End Driving

Abstract

Candidate-based driving planners must select one trajectory from many plausible alternatives. A candidate's driving score can guide this choice, but does not establish its value as a training target. We present RethinkDrive, a lightweight refinement procedure and controlled study of this distinction. With the planner and candidate pool frozen, we compare metric-derived targets with targets anchored to recorded human trajectories, which are used only in training. We evaluate learned rankings before and after combining their scores with the planner's. On NAVSIM, reference supervision improves a diffusion planner, mainly by making successive plans more temporally consistent, with some loss of progress and drivable-area compliance. The gain persists with disjoint training and validation logs and without score fusion; two validation-selected rules do not reproduce it. Fusion even reverses the comparison between a hard reference target and a validation-selected soft target on this planner. On a vocabulary-based planner, it reduces large differences between learned rankings to similar final scores, without a resolved improvement over the baseline. These findings show that supervision quality must be assessed through both the learned ranking and its use in the deployed selection rule.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.