acceptodds
Under review as a conference paper at ICLR 2027

AlterVLA: A Vision-Language-Action Model with Latent Temporal Change Prediction for Autonomous Driving

Abstract

Anticipating scene evolution can improve autonomous driving, but predicting complete future representations may devote modeling capacity to visual content already available in current observations. We introduce AlterVLA, a vision-language-action framework that uses compact, multi-horizon change tokens to focus foresight on scene evolution and provide complementary information for trajectory planning. Learning these representations in the pretrained model's own visual space avoids reliance on an external visual teacher. Predicted changes serve as direct planning conditions within a shared backbone, enabling continuous trajectory generation through a lightweight head without requiring future observations at inference. We further fine-tune AlterVLA with reinforcement learning to optimize safety, progress, and comfort beyond trajectory imitation. Accordingly, a four-stage training pipeline is developed to bridge future-supervised foresight and deployable planning. The planning interface is first learned with oracle changes and then preserved when predicted changes are introduced, before the policy is refined through reinforcement learning. Using a single front-view camera, AlterVLA achieves 88.53 PDMS on NAVSIM v1 and 87.82 EPDMS on NAVSIM v2 under supervised fine-tuning. With reinforcement fine-tuning, it achieves 90.94 PDMS on NAVSIM v1. Ablations show that temporal-change conditioning outperforms full-state prediction and the no-change baseline in PDMS. https://anonymous.4open.science/r/Alter_vla-369D/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.