AlterVLA: A Vision-Language-Action Model with Latent Temporal Change Prediction for Autonomous Driving
Abstract
Anticipating scene evolution can improve autonomous driving, but predicting complete future representations may devote modeling capacity to visual content already available in current observations. We introduce AlterVLA, a vision-language-action framework that uses compact, multi-horizon change tokens to focus foresight on scene evolution and provide complementary information for trajectory planning. Learning these representations in the pretrained model's own visual space avoids reliance on an external visual teacher. Predicted changes serve as direct planning conditions within a shared backbone, enabling continuous trajectory generation through a lightweight head without requiring future observations at inference. We further fine-tune AlterVLA with reinforcement learning to optimize safety, progress, and comfort beyond trajectory imitation. Accordingly, a four-stage training pipeline is developed to bridge future-supervised foresight and deployable planning. The planning interface is first learned with oracle changes and then preserved when predicted changes are introduced, before the policy is refined through reinforcement learning. Using a single front-view camera, AlterVLA achieves 88.53 PDMS on NAVSIM v1 and 87.82 EPDMS on NAVSIM v2 under supervised fine-tuning. With reinforcement fine-tuning, it achieves 90.94 PDMS on NAVSIM v1. Ablations show that temporal-change conditioning outperforms full-state prediction and the no-change baseline in PDMS. https://anonymous.4open.science/r/Alter_vla-369D/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.