Cartender: Verifiable Language-Conditioned Driving by Learning the Drives Not Taken
Abstract
Every driving scene is logged exactly once, yet driving is inherently one-to-many. The scene decides which behaviors are admissible and the driver's intent which one to take: vision alone cannot choose among them, and language alone cannot tell what the scene permits. But the log keeps only the drive that was taken, and its instruction merely describes it. Language-conditioned driving policies are reported to follow instructions, yet on such data this cannot be established: a policy that ignores the instruction and one that obeys it regardless of the scene fit the logs equally well, and neither is distinguishable from one that weighs both. Trained only on such logs, driving vision–language–action models (VLAs) and world–action models (WAMs) receive no direct supervision on actions that violate their instructions or instructions the scene rules out. We present Cartender, which learns from the drives not taken. To construct matched and mismatched instruction–action pairs from real recordings without a simulator, we propose Counterfactual Branching, which treats each logged trajectory as one branch among several candidate futures: a vision–language model proposes the others and the instruction that selects each, and instructions the scene cannot support become negative labels. Cartender then weighs both explicitly: a learned scene–language compatibility regulates language conditioning during trajectory generation, and a reward-trained language residual re-ranks the candidates over a scene-based score. Branches serve training; instruction following is evaluated by changing instructions under identical recorded inputs, using behavior criteria rather than generated trajectory targets. On WOD-E2E, the best of four proposals reaches an RFS of 8.642 across 3,603 feasible instruction conditions. Across all 8,974 admitted proposals, mean RFS is 8.203 versus 7.906 for their matched logged trajectories, and 67% score higher than their logged counterpart. Relative to matched-only imitation with the same backbone, Cartender more than doubles the rate of fulfilling both instructions of a same-scene pair (57.00% vs. 23.67%) while RFS rises from 6.32 to 7.10; its gate cuts execution of infeasible instructions from 53.36% to 14.18%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.