Towards Self-Improving Autonomous Driving through Imagined Interaction
Abstract
Vision-Language-Action (VLA) models have shown strong potential for autonomous driving, yet their learning remains largely confined to fixed offline demonstrations. Such offline training provides little supervision from the states and consequences induced by the policy’s own decisions, limiting continual self-improvement. We present InteractVLA, a framework for self-improving autonomous driving through imagined interaction. Rather than using predicted futures only for reasoning or action evaluation, InteractVLA converts action-conditioned future transitions into multi-step policy-induced experience for policy optimization. To enable efficient imagined interaction, we introduce an Asymmetric Action-Transition Dual Transformer that jointly performs action generation and transition modeling within a single forward pass, while compact transition tokens capture action-induced changes without reconstructing complete future observations. We further propose interaction-aware feedback to maintain reliable supervision as imagined states deviate from logged traffic, together with an interaction-specific return-to-go objective for temporal credit assignment under GRPO. Experiments on NAVSIM and Bench2Drive demonstrate strong planning performance and consistent gains under successive interaction. InteractVLA achieves 93.2 PDMS on NAVSIM and a 91.13 Driving Score on Bench2Drive, and reaches 88.3 PDMS under 5-second continuation on NAVSIM. These results demonstrate the effectiveness of imagined interaction for self-improving autonomous driving.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.