TrackForesight: Track-based Visual Foresight for Automotive VLA
Abstract
Autoregressive vision-language-action (VLA) models have recently attracted growing interest in autonomous driving, owing to their inherited reasoning capabilities, strong pretrained initializations, and well-established architectures. Despite the elegance of this design, standard end-to-end training asks the model to predict sparse action targets directly from raw sensor inputs, which is a difficult supervision problem that leaves little direct supervision for scene structure and agent behavior. We propose TrackForesight, which guides the policy by tracking driving relevant objects in the context via a point tracker, and training the model to predict their future positions as visual chain-of-thought before emitting the final trajectory. We compare point tracks against dense RGB and semantic foresight modalities and show that they offer an excellent balance between predictability, richness, and efficiency, preserving short-horizon motion with far fewer intermediate tokens. Relative to standard action-only training, After GRPO, TrackForesight improves PDMS on NAVSIM navtest by , reaching a state-of-the-art among VLA-based planners on the NAVSIM v1 benchmark, with only generated foresight tokens and no additional training stage. The results suggest that effective foresight for driving VLAs hinges on the information content of the prediction target, the training balance between foresight and action objectives, and the cost of generating intermediate tokens before action.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.