Restoring Spatial Distance Feedback in Token-Native Driving VLAs
Abstract
Discretizing continuous trajectory waypoints directly into a Vision Language Action (VLA) model's vocabulary allows it to treat motor control as standard autoregressive generation. This unified formulation ensures that policy execution directly inherits the pre-training capabilities and scaling law benefits of the underlying Vision Language Model (VLM), bypassing the need for specialized continuous regression heads. However, optimizing these tokenized action spaces with standard categorical cross-entropy imposes an isometric loss landscape that penalizes all incorrect tokens uniformly, regardless of their physical error in metric space. We demonstrate that this loss geometry causes closed-loop policy failure in models trained with discrete K-means tokens (K-DISK) despite high action reconstruction accuracy. We address this bottleneck through two core contributions: 1) adapting Ordered Action Tokenization (OAT) from robotics to autonomous driving, where continuous trajectory proximity is mapped to shared sequence prefixes, and 2) Coarse Trajectory Anchor Regularization (CTAR), a differentiable metric loss that backpropagates smooth- trajectory errors directly through the first coarse token via a frozen decoder. This provides the explicit metric feedback missing from pure cross-entropy. Evaluated on CARLA Bench2Drive and on NAVSIM v2, our method yields significant gains over standard K-DISK (+33.57 DS, +0.0979 EPDMS) and matches the performance continuous action heads (+1.55 DS, +0.0425 EPDMS), demonstrating that VLAs can achieve closed-loop fidelity without departing from native token architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.